The ISEF Grand Award criteria give 15 points to design and methodology — explicitly for whether variables and controls are defined, appropriate and complete — and 20 more to execution, including reproducibility and whether sufficient data was collected. Almost all of that is decided before your first measurement. Design is the cheapest phase of an ISEF project and the one students rush hardest.
Design happens off the clock — and that is a structural gift
The ISEF Rules FAQ defines the project start date precisely: “the start date of your project is when you begin to collect data for your experiment. The literature review and the design of your study will occur prior to your start date.” Only data collection consumes the 12-month research window.
Read commercially, that sentence says: thinking is free, measuring is expensive. A student who spends eight weeks reading, designing, and piloting before the clock starts has lost nothing. A student who starts collecting on week one because it feels like progress has spent window on an experiment they will redesign anyway.
There is a second reason the pre-start period matters. The approval architecture also lives there. The adult sponsor signs on the date they first review the project plan before the experiment begins; risk assessment must be completed and signed prior to student experimentation; and where a project involves human participants, vertebrate animals or biological agents, the relevant review and signature must be dated before any experimentation takes place. Approvals are not paperwork that follows your design — they are a review of your design, and a reviewer who cannot see your variables and controls cannot approve it.
Which produces the practical rule: your design must be legible to a stranger before you own any data. If it is not, you are not ready to start, whatever the calendar says.
What “variables and controls defined, appropriate and complete” actually asks for
That single phrase in the judging criteria is doing four jobs. Students typically answer one of them and assume they have answered all four.
- The independent variable — the one thing you deliberately change, and the specific levels you change it to. “Temperature” is not an independent variable; “15 °C, 25 °C and 35 °C” is. Named levels are what make a design reviewable.
- The dependent variable — what you measure, in what unit, with what instrument, at what moment. Most weak designs fail here rather than at the independent variable: the outcome is named but the measurement procedure is not, so nobody can tell how much of the variation is the phenomenon and how much is the measuring.
- Controlled variables — everything held constant across conditions. This is a list, and it should be written down as one. Judges ask for it because the length and specificity of that list is a fast proxy for how carefully someone has thought.
- The control condition — the comparison group that tells you what happens when you do nothing, or do the standard thing. “Controlled variables” and “a control group” are different objects, and conflating them is the single most common vocabulary error we hear at practice interviews.
A useful self-test before your start date: write each of the four on a separate line and hand the page to someone outside your field. If they can restate what you will change, what you will measure, what you are holding still, and what you are comparing against — you have a design. If they ask a clarifying question, that is the same question a judge will ask, and answering it now costs nothing. The interview logic behind those questions is unpacked in what ISEF judges look for at the booth.

Trials, replicates and what “sufficient data” means when no number is given
Under execution, the criteria ask for systematic data collection and analysis, reproducibility of results, and sufficient data collected to support interpretation and conclusions. Note what is absent: the published criteria we checked state a standard, not a minimum sample size. There is no magic n to hit.
“Sufficient” is therefore relative to your claim. A large effect measured with a precise instrument needs fewer observations than a small effect measured noisily. The design question is not “how many should I do” but “how many do I need before this difference stops being explainable by noise” — which is a question you can answer roughly during a pilot run, before the clock starts.
The distinction that separates a solid design from a fragile one is independent replicates versus repeated measurements. Measuring the same plant five times gives you five numbers about one plant. Measuring five plants once gives you five numbers about the treatment. Both have n = 5 in a spreadsheet; only one supports a claim about the treatment.

What a control looks like in each project type
“Where is your control?” is asked of every project, but the correct answer differs by type — and the engineering criteria are worded differently from the science criteria in exactly this area.
| Project type | What the rubric emphasises | The control question | A weak answer, and a stronger one |
|---|---|---|---|
| Scientific research | “Variables and controls defined, appropriate and complete” | What happens under no treatment, or the standard treatment? | Weak: “I kept conditions the same.” Stronger: a named untreated group run alongside, same day, same operator |
| Engineering | “Prototype has been tested in multiple conditions/trials”; alternatives explored; criteria and constraints defined | Compared against what — the existing solution, or your own earlier version? | Weak: “It works.” Stronger: performance against stated design criteria, plus a baseline device or prior iteration, across several conditions |
| Computational / modelling | Systematic analysis and reproducibility of results | What is the baseline your method must beat, and on what held-out data? | Weak: a single accuracy figure. Stronger: a simple baseline, an identical evaluation split, and repeated runs with different seeds |
The recurring failure across all three is a comparison that is real but never stated. Students frequently do run a sensible baseline and then leave it out of the board and the interview because it felt like scaffolding rather than a result. If your evidence rests on a contrast, the thing you contrasted against belongs in the design section — explicitly, with its own n.
The confounds that get found first
Confounds are variables that move together with your treatment and offer an alternative explanation for your result. Experienced judges probe them in a predictable order, because a small number of them account for most compromised high-school datasets.
- Time. All control samples ran in March, all treated samples in April. Anything that changed between those months is now inside your effect.
- Location. Treatment trays sat nearer the window, the heater, or the door. Position is a treatment nobody assigned.
- Batch and reagent lot. A new bottle, a new supplier, a recalibrated instrument mid-experiment. Record lot numbers; they are free to record and impossible to reconstruct.
- Operator and expectation. One person measured controls, another measured treatments; or the same person knew which was which while scoring a subjective outcome. Where the outcome involves any judgment, blinding the scorer is usually cheap.
- Run order. Conditions were run in a fixed sequence, so any drift in the apparatus maps directly onto condition. Randomising order costs nothing and removes an entire class of objection.
Every one of these is fixable at design time and unfixable at analysis time. That asymmetry is why the pre-start weeks are worth more than they feel like they are worth. It is also why a research question that cannot be operationalised into a clean comparison is not yet a viable project — a point worth applying when choosing an ISEF research topic in the first place.
A one-hour design review before your start date
Run this with your adult sponsor or mentor before any approval signature is dated, and treat any hesitation as a finding.
- State the claim you hope to make in one sentence, in the form “X changes Y by Z under conditions C.” Then ask what design would be needed to earn that sentence. Most students discover their planned design earns a weaker sentence than the one they wanted.
- Write the four variable lines — independent with named levels, dependent with unit and instrument, controlled as an explicit list, control condition as a named group.
- Draw the matrix of conditions by replicates and count the cells. If the total is beyond your time, equipment or site access, cut a level rather than cutting replicates.
- Pilot one full cycle of a single cell, end to end, and time it. Pilot data is the cheapest information you will ever buy, and it tells you whether your instrument resolves the effect at all.
- Name your own three biggest confounds and write the mitigation next to each. Bringing this list to the interview yourself is far stronger than having a judge produce it.
- Freeze the protocol, then get it approved. Once the fair has approved a methodology, changing method mid-season is a separate conversation with your fair — and if you plan to keep collecting between your regional fair and the finals, the rules expect additional data to use the same previously approved methodology. The wider qualification structure is set out in every path to the ISEF finals.
Across the projects Embark has coached — per Embark, more than 750 competition awards overall, with coaching spanning the 22 ISEF categories — the pattern that holds is unglamorous: the projects that survive judging are usually not the ones with the most impressive equipment, but the ones whose comparison was decided, written down and reviewed before the first measurement. Design is the part of research a student can genuinely own without a university lab, and it is weighted accordingly.
What is the difference between controlled variables and a control group?
Controlled variables are held constant across all conditions. A control group is a separate comparison condition given no treatment.
How many trials does ISEF require?
The criteria ask for sufficient data to support interpretation and conclusions rather than a set number. Confirm the rubric on societyforscience.org.
Does designing my study use up the 12-month research window?
No. Per the ISEF Rules FAQ the start date is when data collection begins; literature review and study design occur before it.
Do engineering projects need a control?
They need comparison: the criteria ask that a prototype be tested in multiple conditions or trials against stated design criteria.
Work with Embark
The highest-leverage hour in an ISEF project is the design review before the start date — the point where a confound costs nothing to remove. If you have a question but not yet a matrix, that is exactly the right moment to talk.
Embark is an independent research-coaching organisation, the international competition team of Youfang Education. We are not affiliated with, endorsed by, or sponsored by the Society for Science or Regeneron ISEF. Any results cited reflect Embark's own published record (per Embark). Judging criteria, point weights, research-window definitions and approval requirements change between cycles — confirm all details on societyforscience.org and with your affiliated fair. Factual errors are corrected within 7 working days of notice.