Difference
How a research agent differs from a chat assistant
A chat assistant answers one prompt at a time and waits for your next instruction. A research agent plans its own steps from a goal and adjusts later steps based on intermediate results.
It plans before it acts
Ask it to map the state of a field and it first drafts keywords, source scope and the fields to extract, then searches and organizes. Review that plan: if the direction is wrong, every later step follows it.
It iterates
When the first pass misses, you adjust keywords, source scope or experiment parameters through feedback instead of restating the whole problem.
It helps most across disciplines
Interdisciplinary work spans several vocabularies. The agent can expand the common terms in each field at once, and the researcher filters the results, which beats trying terms database by database.
It can act, not just summarize
An agent with an execution environment can download data, write and run code, and save figures and logs. When a question needs new evidence, it can produce that evidence rather than paraphrase what others did.
Task breakdown
Inputs, outputs and checks for five research tasks
Each task you hand to an agent needs a clear input and an output you can inspect. The check column is the acceptance bar; if a step fails it, do not move on.
| Task | Input | Output | Check |
|---|---|---|---|
| Search | Question, date range, publication types, exclusions | Queries, candidate papers, inclusion reasons | Titles, authors, venues and links open and match |
| Reading | Full texts and the questions to answer | Methods, samples, results, limitations | Each claim maps to a page or passage in the source |
| Hypotheses | Confirmed papers and the research gap | Testable hypotheses with their rationale | Rationale comes from verified papers, and data or experiments could refute it |
| Code and experiments | Hypothesis, data sources, evaluation metrics | Code, data, logs, figures and results | A rerun gives consistent results; parameters and data versions are recorded |
| Writing and revision | Confirmed material, section purpose, content to keep | Cited draft and change notes | Each citation supports its claim; facts and numbers are unchanged |
Method
A research agent workflow you can follow
- 01
Define the task boundary
State four things at once: the research question, allowed sources, deliverable format and stopping condition. For example: "Test whether method A beats method B on a given public dataset, using only peer-reviewed papers and official data. Deliver a comparison table, the code and a one-page summary, and stop after three controlled runs for my review."
- 02
Search first, then read closely
Have the agent generate queries and group candidate papers by theme. Once you confirm the inclusion list, extract each paper's question, method, sample, findings and limitations, and compare experimental setups across papers.
- 03
Derive hypotheses from the gap
Ask the agent to tie each hypothesis to the papers that support it and to state what result would refute it. A hypothesis without a refutation condition is usually not specific enough.
- 04
Write code and run experiments
Start with a small run to confirm data loading, metric calculation and output format, then scale up. Keep code, dependencies, data versions and logs next to the results.
- 05
Let results set the next round
Agree in the first step on what happens if results support, refute or fail to settle the hypothesis. After reviewing the results, the researcher decides whether to iterate, change direction or write up a report.
Scientify
Hand a research goal to Scientify and let it keep going
Scientify is a scientific agent that runs in an isolated cloud computer. Give it a research goal and it keeps refining its search strategy, tries to obtain full-text PDFs, forms hypotheses, and plans and runs experiments, then decides the next round from the results.
Papers, code, dependencies, data, logs, figures and results live in one workspace, so the work can keep iterating and the code is reproducible. The task keeps running after you close your laptop; you can check progress and give feedback from your phone or any device, and end up with code, data and a report.
- The core agent is open source with 2k+ GitHub stars, and the related paper was accepted at ICML 2026
- Connects to compute and licensed scientific software to run chemistry, materials and mechanics calculations and analyze the results
- Conversation history and workspace files stay in the isolated cloud computer; Scientify servers do not store this research data
- New users get a free cloud computer and $5 in model credit, and the same model usage costs about 30% of standard API pricing
Pitfalls
Common mistakes and an acceptance checklist
- Goals that are too broad: "research large language models" has no stopping condition, so the agent just keeps expanding. Narrow it to a question results can answer.
- Reading conclusions, not the process: a fluent report is not sound evidence. Spot-check citations and logs before reading the summary.
- Citations never traced back: if a search result will not open or a summary cannot be located in its source, it stays out of the paper.
- Experiments without a trail: results with no code, parameters or data versions cannot be reproduced or checked by reviewers.
- Outsourcing judgment: the research question, evidence quality, interpretation and academic responsibility still belong to the researcher.
References
- Scientify open-source agent (GitHub) — core agent source code
- Scientify paper (arXiv) — accepted at ICML 2026