AI Agents for Scattering Analysis
I have been adding AI agents to the scattering software I build. They share one design: the agent works on top of a tested scientific core and calls the same code the buttons call; refined values and metrics come from that code, not from the model.
One pattern
Six choices that run through the tools.
A tested core first
The science is deterministic, unit-tested code (1,700+ tests in MATERIA alone). The agent is a new way in, not a new engine.
Numbers come from code
The model chooses what to run and explains the result. Refined values and metrics come from tested code, never from the model.
Guardrails in code
Rules a prompt could forget are enforced in code, from correlation limits to a hard-limit veto.
Evals from real failures
A failure on real data becomes a test, so the fix stays fixed: MATERIA's ten eval scenarios replay in CI.
Local or cloud models
The agents run on a local model (Ollama, LM Studio) or a cloud one, so unpublished data need not leave the machine.
MCP hand-offs
MATERIA and NEXPLAN are also MCP tool servers, and NEXPLAN writes inputs in the formats MATERIA, NEBULA3D and the NeXus Viewer read.
What each agent does
Four scattering tools and one materials-screening experiment.
An Agent beside a powder or PDF fit, on Claude or a local model, that works through the page's own controls. You approve each change or let it run in Auto; every change can be undone, and the engine, not the model, computes the refined values. An MCP server opens the same core to other agents.
An agent beside every page of the 3D-ΔPDF console, on a local or cloud model. It grades each reduction stage from unit-tested metrics, checks the ΔPDF against the crystal's symmetry, and tunes the pipeline from a checked catalog; a trial past a hard limit cannot win.
A chatbox on every page, on a local or cloud model, that answers questions about a reverse Monte Carlo run by calling the Workbench's own analyses; deterministic checks read each result. The Workbench is live; the Copilot is not deployed yet.
My personal SNS experiment planner serves its calculations to agents (on a development branch), with hand-off tools that write the next tool's inputs: an instrument file for MATERIA, symmetry operations and Bragg positions for NEBULA3D and the NeXus Viewer.
An agent that proposes candidate compositions, screens them with physics-grounded surrogates and iterates, on local models by default. Compared with non-LLM baselines under the same relaxation cap, not matched total compute.
Case studies
From validation rounds on real data; each round fixed what it exposed.
PbSO₄: why is the fit poor?
wR 12% → 3.7%
On the GSAS-II tutorial data, the Agent's fit diagnosis found the missing peak asymmetry.
Fluorapatite: is the cell right?
1 part in 10⁵
Once the cell check handled lab data (the Kα₂ doublet, a zero shift), its cell agreed with GSAS-II's refined cell to 1 part in 10⁵.
Cr₂WO₆: do the two cations differ?
P = 0.90
The Agent's Bayesian check gave B(W) > B(Cr) a probability of only 0.90, so one tied B is the defensible model.
Does the ΔPDF keep the crystal's symmetry?
12.6% RMS → agree to rounding
A symmetry check exposed a pipeline defect: six-fold partners differed by 12.6% RMS. Now they agree to rounding, a fix to the pipeline, not an agent result.
Does the AI Copilot hold up on a real model?
5 of 6 and 4 of 6 fully right
Six questions on the demo run: every tool call was valid, and the runs exposed four bugs, now fixed. Both models rated every answer “achieved”, so trust the checks, not the verdict.
Limits
What this work has not shown yet.
- No real-model eval pass rates yet. MATERIA's eval scenarios replay in CI with a scripted model; the AI Copilot's real-model test is a six-question spot check.
- A handful of datasets. The validation rounds cover a few real datasets, not a benchmark.
- Not all released. MATERIA's newest Agent work (its three case studies), the AI Copilot and NEXPLAN's MCP tools are on development branches, not yet in the live apps. Athanor is exploratory.
- Solo work. Personal open-source work by one developer; not peer reviewed.
- A person decides. Guardrails limit what an agent can change, not whether its explanation is right. Check anything you publish against established tools.