Research agent / Workflow

Doing research with an AI agent: setting boundaries and checking the work

A research agent breaks a goal into steps and keeps going: searching, reading, writing code and running experiments. Because it carries on without waiting for you, the opening brief and the acceptance criteria at each step decide whether the output is usable. This guide covers how agents differ from chat assistants, inputs and checks for five task types, a workflow you can follow, and the most common mistakes.

Short answer

To do research with an AI agent, define the research question, allowed sources, deliverable format and stopping condition up front, then let it move through search, reading, hypothesis generation, coding and experiments in order. Every step should be checkable against a source, the code or the logs; the question, the quality of evidence and the interpretation stay with the researcher.

Difference

How a research agent differs from a chat assistant

A chat assistant answers one prompt at a time and waits for your next instruction. A research agent plans its own steps from a goal and adjusts later steps based on intermediate results.

It plans before it acts

Ask it to map the state of a field and it first drafts keywords, source scope and the fields to extract, then searches and organizes. Review that plan: if the direction is wrong, every later step follows it.

It iterates

When the first pass misses, you adjust keywords, source scope or experiment parameters through feedback instead of restating the whole problem.

It helps most across disciplines

Interdisciplinary work spans several vocabularies. The agent can expand the common terms in each field at once, and the researcher filters the results, which beats trying terms database by database.

It can act, not just summarize

An agent with an execution environment can download data, write and run code, and save figures and logs. When a question needs new evidence, it can produce that evidence rather than paraphrase what others did.

Task breakdown

Inputs, outputs and checks for five research tasks

Each task you hand to an agent needs a clear input and an output you can inspect. The check column is the acceptance bar; if a step fails it, do not move on.

TaskInputOutputCheck
SearchQuestion, date range, publication types, exclusionsQueries, candidate papers, inclusion reasonsTitles, authors, venues and links open and match
ReadingFull texts and the questions to answerMethods, samples, results, limitationsEach claim maps to a page or passage in the source
HypothesesConfirmed papers and the research gapTestable hypotheses with their rationaleRationale comes from verified papers, and data or experiments could refute it
Code and experimentsHypothesis, data sources, evaluation metricsCode, data, logs, figures and resultsA rerun gives consistent results; parameters and data versions are recorded
Writing and revisionConfirmed material, section purpose, content to keepCited draft and change notesEach citation supports its claim; facts and numbers are unchanged

Method

A research agent workflow you can follow

  1. 01

    Define the task boundary

    State four things at once: the research question, allowed sources, deliverable format and stopping condition. For example: "Test whether method A beats method B on a given public dataset, using only peer-reviewed papers and official data. Deliver a comparison table, the code and a one-page summary, and stop after three controlled runs for my review."

  2. 02

    Search first, then read closely

    Have the agent generate queries and group candidate papers by theme. Once you confirm the inclusion list, extract each paper's question, method, sample, findings and limitations, and compare experimental setups across papers.

  3. 03

    Derive hypotheses from the gap

    Ask the agent to tie each hypothesis to the papers that support it and to state what result would refute it. A hypothesis without a refutation condition is usually not specific enough.

  4. 04

    Write code and run experiments

    Start with a small run to confirm data loading, metric calculation and output format, then scale up. Keep code, dependencies, data versions and logs next to the results.

  5. 05

    Let results set the next round

    Agree in the first step on what happens if results support, refute or fail to settle the hypothesis. After reviewing the results, the researcher decides whether to iterate, change direction or write up a report.

Scientify

Hand a research goal to Scientify and let it keep going

Scientify is a scientific agent that runs in an isolated cloud computer. Give it a research goal and it keeps refining its search strategy, tries to obtain full-text PDFs, forms hypotheses, and plans and runs experiments, then decides the next round from the results.

Papers, code, dependencies, data, logs, figures and results live in one workspace, so the work can keep iterating and the code is reproducible. The task keeps running after you close your laptop; you can check progress and give feedback from your phone or any device, and end up with code, data and a report.

  • The core agent is open source with 2k+ GitHub stars, and the related paper was accepted at ICML 2026
  • Connects to compute and licensed scientific software to run chemistry, materials and mechanics calculations and analyze the results
  • Conversation history and workspace files stay in the isolated cloud computer; Scientify servers do not store this research data
  • New users get a free cloud computer and $5 in model credit, and the same model usage costs about 30% of standard API pricing

Pitfalls

Common mistakes and an acceptance checklist

  • Goals that are too broad: "research large language models" has no stopping condition, so the agent just keeps expanding. Narrow it to a question results can answer.
  • Reading conclusions, not the process: a fluent report is not sound evidence. Spot-check citations and logs before reading the summary.
  • Citations never traced back: if a search result will not open or a summary cannot be located in its source, it stays out of the paper.
  • Experiments without a trail: results with no code, parameters or data versions cannot be reproduced or checked by reviewers.
  • Outsourcing judgment: the research question, evidence quality, interpretation and academic responsibility still belong to the researcher.

References

FAQ

How do I verify what a research agent produces?

For literature output, confirm the source opens and each summary maps to a specific page or passage. For experiments, confirm the code, parameters, data versions and logs are all there and that a rerun gives consistent results. For key findings, go back to the source and check the subjects, methods and conditions.

Does a research agent work for niche or interdisciplinary topics?

Yes. Have the agent expand keywords across disciplines, then let someone who knows the area filter them. For niche topics, add more databases and citation chasing, and check by hand for missing key terms.

What still has to be done by the researcher?

Choosing the question, selecting material, judging evidence quality, interpreting results and taking academic responsibility. Agents are best at well-bounded work such as searching, organizing, coding, running experiments and drafting.

Does my computer need to stay on for long-running agent tasks?

It depends on where the agent runs. With Scientify, the agent, files and runtime live in an isolated cloud computer, so the task continues after you close your laptop and you can check progress from any device.

Start with one well-bounded research question

Write down your goal, allowed sources and stopping condition, and the agent keeps searching, coding and running experiments in the cloud. New users get a free cloud computer and $5 in model credit.

Start researching