I have been adding AI agents to the scattering software I build. They share one design: the agent works on top of a tested scientific core and calls the same code the buttons call; refined values and metrics come from that code, not from the model.

One pattern

Six choices that run through the tools.

A tested core first

The science is deterministic, unit-tested code (1,700+ tests in MATERIA alone). The agent is a new way in, not a new engine.

Numbers come from code

The model chooses what to run and explains the result. Refined values and metrics come from tested code, never from the model.

Guardrails in code

Rules a prompt could forget are enforced in code, from correlation limits to a hard-limit veto.

Evals from real failures

A failure on real data becomes a test, so the fix stays fixed: MATERIA's ten eval scenarios replay in CI.

Local or cloud models

The agents run on a local model (Ollama, LM Studio) or a cloud one, so unpublished data need not leave the machine.

MCP hand-offs

MATERIA and NEXPLAN are also MCP tool servers, and NEXPLAN writes inputs in the formats MATERIA, NEBULA3D and the NeXus Viewer read.

MATERIA architecture: the web app UI, the in-app Agent, the MCP server and web workers sit on shared parsers and visualization, all calling a pure TypeScript scientific core of thirteen modules with more than 1,700 tests
MATERIA: the page, the in-app Agent, the MCP server and the workers all call one tested core.

What each agent does

Four scattering tools and one materials-screening experiment.

MATERIA In-app Agent · MCP server

An Agent beside a powder or PDF fit, on Claude or a local model, that works through the page's own controls. You approve each change or let it run in Auto; every change can be undone, and the engine, not the model, computes the refined values. An MCP server opens the same core to other agents.

NEBULA3D NEBULA Pilot

An agent beside every page of the 3D-ΔPDF console, on a local or cloud model. It grades each reduction stage from unit-tested metrics, checks the ΔPDF against the crystal's symmetry, and tunes the pipeline from a checked catalog; a trial past a hard limit cannot win.

RMCProfile Workbench AI Copilot (in development)

A chatbox on every page, on a local or cloud model, that answers questions about a reverse Monte Carlo run by calling the Workbench's own analyses; deterministic checks read each result. The Workbench is live; the Copilot is not deployed yet.

NEXPLAN (work in progress) 26 MCP tools

My personal SNS experiment planner serves its calculations to agents (on a development branch), with hand-off tools that write the next tool's inputs: an instrument file for MATERIA, symmetry operations and Bragg positions for NEBULA3D and the NeXus Viewer.

Athanor (exploratory) Closed-loop screening agent

An agent that proposes candidate compositions, screens them with physics-grounded surrogates and iterates, on local models by default. Compared with non-LLM baselines under the same relaxation cap, not matched total compute.

Case studies

From validation rounds on real data; each round fixed what it exposed.

MATERIA · neutron powder (D1A)

PbSO₄: why is the fit poor?

wR 12% → 3.7%

On the GSAS-II tutorial data, the Agent's fit diagnosis found the missing peak asymmetry.

MATERIA · lab X-ray (Cu Kα)

Fluorapatite: is the cell right?

1 part in 10⁵

Once the cell check handled lab data (the Kα₂ doublet, a zero shift), its cell agreed with GSAS-II's refined cell to 1 part in 10⁵.

MATERIA · neutron powder, 150 K

Cr₂WO₆: do the two cations differ?

P = 0.90

The Agent's Bayesian check gave B(W) > B(Cr) a probability of only 0.90, so one tied B is the defensible model.

NEBULA3D · a measured 6/mmm volume

Does the ΔPDF keep the crystal's symmetry?

12.6% RMS → agree to rounding

A symmetry check exposed a pipeline defect: six-fold partners differed by 12.6% RMS. Now they agree to rounding, a fix to the pipeline, not an agent result.

RMCProfile Workbench · two local models

Does the AI Copilot hold up on a real model?

5 of 6 and 4 of 6 fully right

Six questions on the demo run: every tool call was valid, and the runs exposed four bugs, now fixed. Both models rated every answer “achieved”, so trust the checks, not the verdict.

Limits

What this work has not shown yet.

  • No real-model eval pass rates yet. MATERIA's eval scenarios replay in CI with a scripted model; the AI Copilot's real-model test is a six-question spot check.
  • A handful of datasets. The validation rounds cover a few real datasets, not a benchmark.
  • Not all released. MATERIA's newest Agent work (its three case studies), the AI Copilot and NEXPLAN's MCP tools are on development branches, not yet in the live apps. Athanor is exploratory.
  • Solo work. Personal open-source work by one developer; not peer reviewed.
  • A person decides. Guardrails limit what an agent can change, not whether its explanation is right. Check anything you publish against established tools.