Before you build a single slide, you need to know whether your validation is credible. This article walks through a leakage audit that separates defensible machine-learning research from numbers that evaporate under scrutiny.
Youfang Education’s August 1, 2026 briefing attributes to Dr. Wu the observation that finding a meaningful question matters more than chasing fashionable techniques or benchmark accuracy, and that communication of the project matters alongside experiments. This principle guides the following audit: a strong question poorly validated is still weak science. The audit below is developed for this article—it is not a pre-existing Embark protocol, but a framework you can apply immediately.
Why Data Leakage Destroys Credibility
Machine learning projects at science fairs face a special trap. Unlike a chemistry experiment where a failed replication is visible, an overfitted model can look successful until someone examines the data pipeline. Leakage—information from outside the training set influencing model training or evaluation—produces optimism that disappears when the model meets truly new data. The seven patterns below are useful checks for a student research project.
Seven Leakage Patterns to Eliminate
Prediction target and unit of independence undefined. Before touching any split, state exactly what you are predicting and what constitutes an independent instance. If your target is “disease present on a plant,” then the plant—not the photograph, not the leaf patch—is your unit of independence. Confuse these, and your split logic unravels.
Repeated records across partitions. The same person, device, or specimen appears in both training and testing. The model can rely on shared characteristics rather than generalizable patterns.
Near-duplicates or augmented images crossing splits. Five photographs of the same plant, rotated and color-adjusted, scattered across train and test. The model may recognize plant-specific details rather than disease patterns.
Fitted preprocessing or feature selection on all data. Standardizing using global mean and variance, or selecting features based on full-dataset correlations, embeds test information into training.
Target-derived fields unavailable at prediction time. Including a variable computed from the target or from future information that would not exist when the prediction is actually needed.
Future rows informing past predictions. In temporal data, training on 2023-2024 to predict 2022, or using any future-dated information.
Repeated tuning on the test set. Treating the final evaluation set as a development playground, adjusting until numbers satisfy, then presenting it as untouched validation.
The Plant Dataset: A Worked Hypothesis
Consider a concrete scenario. A greenhouse ships 40 plants. You photograph each plant 5 times under varying conditions, yielding 200 images. Your goal: predict disease on a new plant the greenhouse has never seen. You still classify images, but the claim being tested is whether performance transfers to an unseen plant.
Here is where naive splitting fails. A row-wise random split of 200 images can place multiple photographs of Plant 17 in both training and test. The model may learn Plant 17’s specific leaf texture, pot color, or background wall. When “tested,” it appears accurate. Ship it to a new greenhouse, and the model may fail because it relied on plant-specific shortcuts rather than disease mechanisms.
The correct split is by plant_id, keeping all derived images with their parent. All 5 photographs of Plant 17 travel together into one partition. This blocks that route to specimen memorization, although other shortcuts can remain.

Designing Your Split: An Illustrative Assignment
The table below shows one possible assignment of 40 plants to partitions. It is explicitly illustrative—not a universal ratio. Your actual split depends on total sample size, expected effect size, and whether you need multiple development folds.
| Partition | Plant IDs (example) | Plants | Images | Purpose |
|---|---|---|---|---|
| Training | 1–24 | 24 | 120 | Model fitting, hyperparameter search via cross-validation |
| Development (validation) | 25–32 | 8 | 40 | Model selection, early stopping, architecture decisions |
| Final test | 33–40 | 8 | 40 | Single evaluation after all decisions finalized |
Every plant appears exactly once. The 200 original images are partitioned by plant_id, not by individual image. Augmented derivatives generated after this assignment remain with their parent plant’s partition.
What Pipelines Protect—and What They Don’t
Scikit-learn’s pipeline mechanism, as documented in their common pitfalls guide, ensures that preprocessing is fit on training folds and only transform applied to held-out data. This prevents fitted preprocessing leakage when the whole pipeline is fitted on the training data within each cross-validation fold. Similarly, GroupKFold keeps an entire group in either the training portion or the validation portion of a fold, never both, which directly supports the plant_id split above. TimeSeriesSplit addresses chronological ordering for temporal prediction tasks.
However, these tools do not fix subject overlap if you have already violated the group boundary. They do not remove a target-leaking input column that you mistakenly included. They do not prevent you from peeking at final-test results and iterating. Pipeline protects learned preprocessing boundaries; it does not substitute for correct split design and research discipline.

Generalization to a New Greenhouse
The plant_id split solves one problem but raises another. Your 40 plants came from Greenhouse A, with its specific humidity, lighting, and camera setup. Even with perfect group splitting, your model may learn Greenhouse A’s artifacts: shadow patterns, pot colors, background shelving. Generalization to Greenhouse B is a separate scientific claim requiring separate validation.
Be explicit about this limitation. If you only have Greenhouse A data, limit your evaluation claim to the held-out plants in that environment and report the measured performance and uncertainty; this does not establish universal deployment. Propose follow-up: collecting data from Greenhouse B to test environmental robustness. This honesty strengthens, not weakens, your project.
Consistent Baselines and the Final Test
Establish consistent baselines before touching your model. A simple majority-class classifier. A shallow decision tree with default parameters. These anchor your sophisticated model’s gains in context. Without them, a high accuracy figure is meaningless noise.
Use your development set for model choice: architecture, hyperparameters, feature sets. Then evaluate once on the final test. Merely viewing or recomputing the same frozen evaluation again does not by itself invalidate it. However, adaptive model selection, feature engineering, or threshold tuning based on final-test outcomes compromises an unbiased final estimate. Once you have used the final test for such decisions, you must disclose this or seek a new genuinely independent evaluation.
Never present a final test as “untouched” after you have adjusted based on its feedback. This undermines the credibility of your validation.
What to Write in Your Notebook
Your research notebook should contain:
- Dataset version, license, and source: Where obtained, when downloaded, any preprocessing applied before your analysis began.
- IDs and split manifest: Exact list of which plant_ids (or equivalent) belong to which partition, with random seed and code version that generated it.
- Transformations fitted where: For each preprocessing step, note whether it was fit on training data only and how applied to other partitions.
- Seed and code versions: Software versions, random seeds, any hardware-specific notes.
- Hypothesis and model-selection record: What you tried, why you tried it, what development metrics guided each choice, and when you stopped.
This documentation is not bureaucratic overhead. It is the evidence that your validation is defensible.
Questions a Mentor Can Ask
A mentor reviewing your project might ask:
- “If I gave you a photograph of a plant you have never seen, not just a new photograph of a familiar plant, how would your model perform?”
- “Show me the exact rows that went into your final test, and prove no derived images from the same source appear in training.”
- “What would change if this model were deployed in a greenhouse with different lighting?”
- “When did you last look at your final test results, and what decisions followed?”
- “What simpler model did you compare against, and what does the gap tell us?”
These are not official judging questions. They are the kind of probing that separates rehearsed presentation from rigorous research.
FAQ
Can I use data augmentation if I split by plant_id?
Yes, but augmented versions must stay with their parent plant. If Plant 17 is in training, all rotations, flips, and color jitters of Plant 17 also belong in training. Never augment into another partition.
What if I only have 20 plants total?
Small sample size makes group splitting harder but more essential. With only 20 plants, there is no universal count for a held-out final test; every withheld plant reduces training data substantially. Consider group-aware development validation (for example, leave-one-plant-out or small-group cross-validation) and report broad uncertainty intervals. Be transparent that generalization evidence is limited. Statistical methods for small samples may help.
Is it ever acceptable to look at final test results multiple times?
Merely viewing or recomputing the same frozen evaluation again does not by itself invalidate it. However, adaptive model selection, feature engineering, or threshold tuning based on final-test outcomes compromises an unbiased final estimate. Once you have used the final test for such decisions, you must disclose this or seek a new genuinely independent evaluation.
Building Defensible Research
Data leakage is not a minor technicality. It is the difference between a model that works in demonstration and one that works in reality. For ISEF-bound students, the audit above provides a framework: define your unit of independence, split before you explore, protect your final test, and document everything. Your presentation will be stronger not because the numbers are higher, but because they mean something.
This independent coaching article is not affiliated with, endorsed by, or sponsored by Society for Science or Regeneron ISEF. This article provides methodological guidance and does not establish regulatory approval or official endorsement.
Sources: Scikit-learn Common Pitfalls; Scikit-learn Cross-Validation; Youfang Education, “ISEF 2026-2027 Season Preparation Guide,” August 1, 2026. Related: ML projects, research plan/notebook, data analysis.