EVOLUTION
Start from a new archive, literature, or human-annotation input; generate, challenge, sanity-check, and re-rank ideas. Without new input, the loop legally stops at STAGNATION rather than spending tokens remixing old material.
EVIDENCE-DRIVEN IDEA EVOLUTION
XMM Proposal Agent is an XMM-Newton research framework. It turns new nearby-galaxy X-ray archive inputs, idea generation and attack, feasibility calculations, and TAC-style non-compensatory judgment into a traceable long-lived evolution line.
The core is not one oversized prompt. It is a set of explicit state boundaries: stable role specifications, accumulating scientific memory, and disposable but auditable execution traces.
Role specifications. Changed rarely; they say who may do what and what they must not do.
Reusable archive coverage, idea ledger, sources, and trust records.
Materials, attacks, rebuttals, calculations, and handoffs from every round.
≠The critical separation: shared memory is not silently rewritten by one failed run, and a run cannot redefine the role rules that it relies on. tools/ is a separate calculators layer for reproducible feasibility and provenance commands; it does not decide scientific state.
This conceptual poster treats the system as an observatory made of distinct stations. Materials, archive facts, ideas, calculations, and drafts may travel only along a trace-preserving rail; no one role skips every check.
Cyan and lime rails depict the traceable flow of evidence packets, shared memory, and did / found / next handoffs.
The red station is the R4 red-team: it stress-tests the blueprint rather than finding better wording for it.
The branching gate and round table represent R5 portfolio decisions and the R8 TAC-sim panel.
This is an architecture concept, not an execution record. Concrete scientific state, proposal lifecycle, and source-verified facts remain in the repository's canonical records.
This prevents the reversed workflow of choosing a topic first and searching for reasons later. The system maintains a portfolio repeatedly filtered by evidence and critique; only a mature idea enters a formal proposal run.
Start from a new archive, literature, or human-annotation input; generate, challenge, sanity-check, and re-rank ideas. Without new input, the loop legally stops at STAGNATION rather than spending tokens remixing old material.
Begin only with a proposal-ready idea. An ObsID-level archive check and the R6-full feasibility chain come first; when numbers do not close, R5 may only narrow or replace the objective. R8 TAC-sim then grades the draft against an explicit rubric.
The framework separates an attractive idea from an idea that has survived review. Attacks, eliminations, and unresolved empirical questions are not concealed; they become explicit inputs to the next cycle or a future proposal's deepening plan.
A new round must consume new external information, a verified archive record, or an explicit human annotation. Stopping is an informative state, not a failure.
R2 proposer, R4 red-team, and R5 judge are distinct fresh shifts. Independence is more than a new label: it prevents end-to-end attachment.
Logical or physical fatal problems, salvageable framing failures, and questions that require new data enter different downstream paths.
R5: any single poor axis can block promotion. R8: only when two or more reviewers assign poor to the same dimension is a draft returned for revision.
Roles are not ceremony. They separate retrieval, evidence verification, scientific imagination, adversarial review, numerical closure, and writing so that an error has an identifiable source and repair point.
Discovers public candidate material; it contributes discovered leads, not source-verified truth.
Turns multi-observatory archive facts into reusable coverage and opportunity records.
Performs source and ObsID-level archive work for one question, preventing a duplication narrative.
Builds baseline ideas and forks from verified opportunities and questions, not invented evidence.
Gathers evidence for and against an idea, preserving unknowns and the data needed to resolve them.
Attacks duplication, observatory fit, decisiveness, physics, and cost as an XMM TAC referee would.
Manages the portfolio; it kills, parks, forks, or promotes under explicit axes and min-gates rather than average impressions.
Turns model → count rate → pileup → background → exposure into an inspectable technical chain.
Organizes the closed evidence-and-feasibility record into an XMM AO structured draft.
Uses different reviewers and the science merit / appropriate use of XMM / observational realizability / completeness & clarity rubric to grade the draft and route concrete revisions.
Turns unclear instrument and method concepts into learnable, testable explanations so the human remains able to judge.
A public page may explain the design, but it must not package a design aspiration as a validated scientific conclusion. Project mechanisms have Trusted, Provisional, and Questioned states. Proposal lifecycle, instrument facts, and any specific scientific result remain authoritative only in their respective canonical records.
For an XMM proposal, scarce value is not wording. It is a scientific opportunity that remains after archive checks, counterarguments, feasibility, and panel judgment.
Return to the framework