Explain relationships
Regression, causal graphs, ANOVA
Definitions, effect sizes, residuals, alternatives
Five-stage workflow
A model is not produced in one leap. Recording each stage makes it possible to locate failures in the definition, data, method, or interpretation.
What must be explained, predicted, optimized, or simulated?
Which mechanisms matter and which details can be omitted?
How will parameters, states, optima, or simulations be obtained?
Is the result reliable beyond one setting?
What did the model answer, and what did it not answer?
01 · Define the problem
Input
Prompt, context, data notes, and real constraints.
Key action
Separate targets, decisions, constraints, evaluation criteria, and scope.
Stage output
A computable problem statement, variable table, and constraint list.
Failure signal
The objective is vague or the metric does not match the real goal.
01
What must be explained, predicted, optimized, or simulated?
Input
Prompt, context, data notes, and real constraints.
Key action
Separate targets, decisions, constraints, evaluation criteria, and scope.
Stage output
A computable problem statement, variable table, and constraint list.
Failure signal
The objective is vague or the metric does not match the real goal.
02
Which mechanisms matter and which details can be omitted?
Input
Problem statement, domain knowledge, data quality, and scale.
Key action
State assumptions, relationships, boundary conditions, and candidate model families.
Stage output
Assumptions, notation, relationship diagram, or equations.
Failure signal
Assumptions cannot be checked, or units and scales are inconsistent.
03
How will parameters, states, optima, or simulations be obtained?
Input
Model, data, initial values, parameter ranges, and compute limits.
Key action
Choose algorithms and tools; record versions, seeds, solver settings, and convergence.
Stage output
Runnable code, estimates, solutions, or simulation results.
Failure signal
Only screenshots remain; code and execution conditions are missing.
04
Is the result reliable beyond one setting?
Input
Results, baselines, held-out data, and alternative specifications.
Key action
Check errors, residuals, convergence, sensitivity, robustness, and edge cases.
Stage output
Validation tables, error plots, sensitivity results, and failure conditions.
Failure signal
Only the best result is shown, without a baseline, error, or counterexample.
05
What did the model answer, and what did it not answer?
Input
Validated results, the original problem, and intended use.
Key action
Connect results to the real task and state limitations, uncertainty, and scope.
Stage output
Traceable figures, conclusions, report, and reproduction notes.
Failure signal
Claims exceed the data or numbers cannot be traced to computations.
Start with the question, then compare assumptions, data needs, and validation evidence.
Regression, causal graphs, ANOVA
Definitions, effect sizes, residuals, alternatives
Time series, classification, regression, ML
Data split, baseline, error distribution, drift checks
Linear, integer, dynamic, heuristic optimization
Objective, constraints, feasibility, optimality or bounds
Differential and difference equations, system dynamics
Initial values, parameter sources, stability, sensitivity
Monte Carlo, queueing, discrete-event simulation
Distributions, seeds, repetitions, confidence intervals
Multi-criteria methods, dimension reduction, clustering
Metric direction, weights, scaling, robustness
Validation gates
Validation asks whether a model can support a claim. Metrics vary, but the checking logic is consistent.
Do the code and equations behave as intended?
Unit, dimension, conservation, and boundary checks
How does the model perform on known and unseen data?
Residuals, holdout error, cross-validation, baselines
Do parameter, sample, or initial-value changes alter the claim?
Sensitivity, robustness, intervals, scenarios
For which objects and ranges does the claim hold?
Limits, failure cases, extrapolation bounds, domain review
If these six fields are unclear, improve the problem definition before searching for algorithms.
No. Simple models can be solved analytically or with spreadsheets. Programming becomes valuable for data processing, estimation, optimization, simulation, and repeatable validation.
Define the question first, then examine data availability and model requirements together. Choosing only from available data can miss the goal; choosing an ideal model may leave no usable evidence.
No. Complexity should match the question, data, and ability to validate. A simpler model with explicit assumptions and stable results is often more useful than an opaque complex model.
Check implementation correctness, comparison with observations or baselines, stability under parameter and sample changes, and the scope in which conclusions apply.
Scientify can organize data, run code, save results, and draft a report in one cloud workspace. You retain control over assumptions and conclusions.