Criteria
Six requirements for a defensible research hypothesis
AI readily produces sentences that are grammatical and sound novel. Checking each candidate against these six requirements filters out most of the ones that are really just opinions.
Grounded in a literature gap
Novelty comes from the gap, not the wording. A hypothesis generated without reference to the papers you have read will not survive the question: what is your evidence?
Feasible with what you have
Can you get the data, is the method mature, are there ethical or compute limits? Where it gets stuck tells you whether the direction belongs in a proposal.
Worth doing
Would closing the gap move the field forward? A thinly studied area may be a real gap, or it may be thin because nobody found it worth the effort.
Stated in testable form
Are the independent variable, dependent variable and expected direction explicit? If you cannot say what you would measure, you have a position, not a hypothesis.
Falsifiable
What data or experimental outcome would disprove it? If there is no answer, it is not a hypothesis yet. Falsifiability is the line between a hypothesis and an opinion.
Specific at the right level
"X affects Y" is too broad. "Under condition Z, X increases Y through mechanism M" gives you something to design an experiment around.
Gaps
Types of literature gaps and the hypotheses each one supports
When AI reads the literature for you, do not just ask what has not been done. Probe one gap type at a time and require every claim to point back to a specific paper.
| Gap type | Signal in the literature | Hypothesis it can become |
|---|---|---|
| Conflicting findings | Similar studies report opposite or inconsistent effects for the same relationship | The difference is driven by a moderator such as sample, conditions or measurement |
| Methodological gap | Existing findings rest mainly on one method, small samples or simulated data | The finding holds, or fails, under a stricter method or a larger dataset |
| Untested setting | The relationship has only been checked in a particular population, material, region or period | The relationship holds in a new setting, or reverses because one condition differs |
| Unexplained mechanism | A correlation is reported repeatedly but the pathway has never been tested | X affects Y through mediator M, and blocking M weakens the effect |
| Limitations authors admit | Future work and uncontrolled factors listed in the discussion section | Controlling for that factor changes the effect size in a predictable way |
Method
Four steps from literature gap to testable hypothesis
- 01
Scope the field with one round of search
Specify the subject, methods and time range, retrieve the core papers and group them by theme. Compare subjects, methods and findings across groups to find where studies contradict each other or leave ground uncovered.
- 02
Work through the gap types one by one
Using the table above, look for conflicting findings, methodological gaps, untested settings, unexplained mechanisms and admitted limitations. Every judgment should point to a specific paper and passage; set aside anything that cannot.
- 03
Generate several candidates, then screen them together
Ask the AI for multiple candidate hypotheses per gap instead of a single answer. Score them against the six requirements and keep the one or two with the clearest support and the most accessible data.
- 04
Write a falsifiable statement and argue against it
Phrase each survivor as "under condition Z, a change in X will move Y in a stated direction" and write the matching null hypothesis. Then have the AI argue the reviewer's side, and rerun searches with new keywords to confirm nobody has already tested the same claim.
Testing
After the hypothesis: testing it with data and experiments
A hypothesis is ultimately judged by evidence. Most AI tools stop once a hypothesis is proposed, but getting data, running the analysis and iterating is where most of the time goes.
- 01
Set decision criteria before looking at data
Write down the primary outcome, the comparison, the statistical test, and what effect size or direction counts as support or refutation. Changing the criteria after seeing the data undermines the conclusion.
- 02
Pick the cheapest test that works
If a public dataset can test it, you may not need a new experiment yet; if simulation or computation can rule a direction out, do that first. Run a small pilot to confirm the pipeline works before scaling up.
- 03
Keep the code, data and logs
Every analysis should be traceable to its data version, code, parameters and output. When results surprise you, you need to tell whether the hypothesis was wrong or the pipeline was.
- 04
Revise or drop the hypothesis based on results
When results do not support it, separate a genuine refutation from weak measurement, sampling or methods. Record the first honestly; only the second calls for a redesigned round.
Scientify
Use Scientify to go from hypothesis to experiment
Scientify is a scientific agent that runs in an isolated cloud computer. Give it a research goal and it organizes its search strategy, tries to obtain full-text PDFs, forms hypotheses, then plans and runs experiments, deciding the next round from the results until it makes research progress.
Papers, code, dependencies, data, logs, figures and results live in one workspace, so each round of testing is reproducible and the next round builds on the last. For hypotheses in chemistry, materials or mechanics, the agent can connect to compute and licensed simulation software to build models, run calculations and analyze the output. The task keeps running after you close your laptop, and you can check progress and give feedback from your phone.
- The core agent is open source, with 2k+ GitHub stars
- The related paper was accepted at ICML 2026
- Conversation history and workspace files stay in the isolated cloud computer; Scientify servers do not store this research data
- New users get a free cloud computer and $5 in model credit
Division of labor
Judgments that stay with the researcher
- Whether the gap is real: check AI-identified gaps against the original papers and search again to confirm nobody has filled them.
- Research value: sparse does not mean important; someone who knows the field has to decide whether it is worth doing.
- Decision criteria: the researcher sets what counts as support or refutation before seeing the data.
- Interpretation: why the results do or do not support the hypothesis, and how far they generalize, is the researcher's call.
References
- Scientify open-source agent (GitHub) — Core agent source code
- Scientify paper (arXiv) — Accepted at ICML 2026