SimAuthor

Harnessing Foundation Models for Persistent Scientific Simulator Authoring

Yishan WangRan PiaoMathias FunkAaqib Saeed

Eindhoven University of Technology

Real recording
search setreference
Root program S(0)
zero-shot
Authored program
after 100 attempts

Overview

A scientific simulator is a program that encodes how mechanisms generate an observable signal. With no exact test of correctness, it has to be revised against a few real recordings. SimAuthor keeps every simulator it writes, chooses what to revise by score and how to revise from a discrepancy report, and reuses edits that worked.

On six biomedical tasks it beats score-only search on all six and reaches the best final score on five. Most of the gain is structural: 111 of 138 audited revisions change program structure, accounting for 86.1% of the signed improvement.

Overview: LLM prior, zero-shot simulator, authoring loop and authored simulator for COPD.
Figure 1. SimAuthor turns LLM prior knowledge into executable simulator code when exact verification is unavailable and feedback comes only from scarce real data.
The SimAuthor authoring harness
Figure 2. A search orchestrator selects simulator candidates based on numerical scores, while the refiner updates them using the scientific blueprint, discrepancy report, and mechanism library.

Method

  1. Evaluate each program against the real recordings, giving a score and a discrepancy report.
  2. Select a parent by score with flat PUCT.
  3. Refine it with one focused edit, guided by its report, a scientific blueprint and recent mechanisms.
  4. Accumulate improving edits as reusable mechanisms.

Explore

Drag the slider to replay a 100-attempt run program by program. Select any program in the tree to see its code change, the feedback it received and the context the refiner was given; the diagram highlights the part of the harness behind each tab.

Harness
Search tree
Score–lineage of the best program

Shown for each task: the seed whose best program appears latest (VSD and COPD seed 1, AS seed 2, AF seed 1, LQT and WPW seed 3).

Results

Search performance

Best-so-far search score over 100 authoring attempts, recomputed from the released runs. Lines show the median-seed run and shading spans three seeds; Sampling is a single pool in fixed proposal order. PUCT score search uses the same selection rule as SimAuthor but gives the refiner only the score, so the gap reflects what the refiner is shown rather than how candidates are chosen. SimAuthor uses about 3.8× the tokens of best-of-100 sampling; a token-matched best-of-380 still falls short on all six tasks.

Generalization

Each test changes one thing relative to search: the recordings, the evaluator, or what the simulator is used for. Every test covers all tasks to which it applies.

New recordings

All six tasks

The search evaluator on held-out recordings: patient-disjoint for VSD and AS, external datasets for the other four. Median of three seeds.

Table

New evaluator

Audio and PPG tasks · ECG search already uses ECGFounder

The program chosen by the search score, re-scored with encoders never used during search: OPERA-CE for audio, PaPaGei for PPG.

Table

Downstream utility

ECG tasks · labelled PTB-XL and Chapman cohorts

AUPRC of a linear probe on frozen ECGFounder features after adding 300 synthetic positives. ID: PTB-XL; OOD: Chapman–Shaoxing. Probe and search share ECGFounder.

Table

Citation

@article{wang2026simauthor,
  title   = {{SimAuthor}: Harnessing Foundation Models for Persistent Scientific Simulator Authoring},
  author  = {Wang, Yishan and Piao, Ran and Funk, Mathias and Saeed, Aaqib},
  journal = {arXiv preprint arXiv:2610.06257},
  year    = {2026}
}

Scope

Search scores measure agreement between simulated and real signal distributions. They do not establish physiological or mechanistic fidelity, and the model can only correct discrepancies the evaluator exposes. All data are public, de-identified research datasets used under their original licences; no recordings are redistributed. We make no diagnostic or clinical claims.