01Overview
OpenAlz is a long-horizon, agentic research platform for Alzheimer's disease (AD). A Master Orchestrator coordinates a swarm of specialised sub-agents that continuously read public biomedical data, generate falsifiable hypotheses, verify their own citations, debate one another, audit for bias and — crucially — learn from every cycle so the next one is better.
OpenAlz is built as a single, transparent research service: one orchestrator, nine specialist agents, an ethical reasoning gateway, a bias detection service, an agent registry and an append-only audit trail — all reading five live public databases and writing to one shared evidence store.
02Quick start
- 1Open the foundry
Go to /app. The Command Center shows live stats and the LLM queue.
- 2Set or launch a goal
Type a research goal (e.g. “TREM2 signalling and tau propagation”) and press Launch run — or create a long-horizon goal under Research Goals and press Pursue.
- 3Watch the agents
The run page streams every agent step; expand any step to see the raw evidence (PMIDs, NCT IDs, targets, pathways).
- 4Read the insights
Each hypothesis shows novelty / confidence / rigor, verified citations, the debate, bias flags and suggested collaborators. Convene further debates from any insight.
- 5Let it evolve
Turn on Autonomous mode in Settings for daily scans, and let self-correction run every N runs. Inspect Strategy Memory to see what it learned.
03How self-evolution works
OpenAlz doesn't just answer — it changes itself. Four feedback mechanisms close the loop:
| mechanism | what happens | effect on future runs |
|---|---|---|
| run debrief | At the end of each run the orchestrator reviews validator scores, debate resolutions, the bias audit and failed agents, and distils three reusable lessons. | The newest 12 lessons are injected into every agent's system prompt. |
| self-correction | Every N completed runs (or on demand) the orchestrator reviews agent success/failure counters, recent failures, score distributions and current memory. | Writes adaptations to memory, retires stale lessons, proposes new goals. |
| agent reflection | Any agent can critique its own recent steps: strengths, weaknesses, a self-rating and concrete adjustments. | Adjustments are injected into that agent's own instructions. |
| daily scan | In autonomous mode the orchestrator wakes daily, reads the last 7 days of PubMed and trial updates, and sets or advances a goal. | The research agenda follows new evidence without a human prompt. |
Scan ─▶ Plan ─▶ Retrieve ─▶ Hypothesise ─▶ Debate ─▶ Audit ─▶ Learn ▲ │ └────────── strategy memory · reflections · goals ◀────────┘
04The agents
| agent | role | grounded in |
|---|---|---|
| Master Orchestrator | Daily scans, goal setting, task decomposition, hypothesis synthesis, debate arbitration, self-correction. | LLM reasoning |
| Literature Bridger | Searches AD literature and an orchestrator-chosen adjacent field; finds mechanistic bridges and gaps. | PubMed E-utilities |
| Biomarker Hunter | Proposes fluid / imaging / genetic / digital biomarkers with evidence strength. | PubMed + Open Targets |
| Drug Screener | Shortlists clinical-stage and repurposable drugs; identifies untargeted genes. | ChEMBL + Open Targets |
| Trial Optimizer | Analyses trial landscape, failure patterns, endpoints and cohort gaps. | ClinicalTrials.gov v2 |
| Pathway Modeler | Enriches the evidence gene set and models pathway crosstalk. | Reactome Analysis Service |
| Data Harmonizer | Unifies entities across agents; flags convergent signals & inconsistencies. | All agent outputs |
| Hypothesis Validator | Verifies citations, raises the strongest objection, scores novelty/confidence/rigor. | PubMed + ClinicalTrials.gov |
| Collaboration Matchmaker | Matches insights to real authors from retrieved papers. | PubMed author records |
| Bias Detection Service | Audits evidence, trial cohorts and hypotheses for bias; powers the Bias Portal. | Trial demographics + text |
05Data sources
| source | used for | endpoint |
|---|---|---|
| PubMed | Literature search (5-year window; 7-day for daily scans), abstracts, authors, citation verification | eutils.ncbi.nlm.nih.gov |
| ClinicalTrials.gov | AD trial landscape, enrollment, age & sex criteria, NCT verification | clinicaltrials.gov/api/v2 |
| Open Targets | Top AD (MONDO_0004975) target–disease associations | api.platform.opentargets.org |
| ChEMBL | Drug indications for Alzheimer's disease | ebi.ac.uk/chembl/api |
| Reactome | Pathway over-representation for the evidence gene set | reactome.org/AnalysisService |
06Methodology & scoring
For each run the orchestrator generates three hypotheses grounded only in identifiers present in the retrieved evidence. The Hypothesis Validator then scores each one:
| score | meaning | notes |
|---|---|---|
| novelty | How far the hypothesis is from prior art and from the foundry's own previous insights | 0–100 |
| confidence | Strength and consistency of the cited evidence | Penalised for unverified citations |
| rigor | Falsifiability and quality of the proposed validation experiment | 0–100 |
| verdict | Final orchestrator ruling after arbitration | accept · revise · reject |
Citation verification: every PMID is checked with PubMed esummary; every NCT ID with the ClinicalTrials.gov API. Each citation is shown with ✓ or ✗ in the UI and report.
07Multi-agent debate
Debates happen inside every run (validator objection → proponent rebuttal → orchestrator resolution) and on demand from the Orchestrator page or any insight. An on-demand debate retrieves fresh PubMed evidence, then runs four turns — proponent (Literature Bridger), skeptic (Hypothesis Validator), rebuttal, counter — before the orchestrator rules supported, contested or refuted with a confidence and next steps. The verdict is attached to the insight's debate history.
08Privacy & ethics
- No patient data. Only public, aggregate sources are queried.
- Ethical LLM gateway. Chat prompts are screened for prompt-injection patterns and PII (emails, phone numbers, SSNs are redacted before reaching the model); flags are shown to the user.
- Full audit trail. Every step start/finish/failure, self-correction and debate resolution is logged.
- Bias detection. Each run ends with a bias audit; anyone can submit text to the Bias Detection Portal.
- Limits. LLMs can be wrong; abstracts are truncated; short gene symbols can match unrelated papers; scores are model judgements, not statistics. Not medical advice.
09Architecture
React (landing · /doc · /app) ──HTTPS──▶ FastAPI /api
├─ orchestrator + 9 agents (asyncio tasks)
├─ priority LLM limiter (chat > research)
├─ autonomy: daily scan · self-correction · debates · reflection
├─ ethical gateway · bias detection · notifications
└─ MongoDB: runs · insights · goals · learnings · audit_logs …
Platform cron (every 15 min) ──▶ /api/cron/tick ──▶ due schedules + autonomous daily scan| component | responsibility |
|---|---|
| orchestrator | Run pipeline, daily scan, self-correction, debate arbitration |
| specialist agents (9) | Literature, biomarkers, chemistry, trials, pathways, harmonisation, validation, matchmaking, bias |
| reasoning gateway | Swappable engines, priority queue (chat before research), prompt-injection & PII screening, streaming |
| bias detection | Bias audit in every run plus the open Bias Detection Portal |
| audit trail | Append-only log of every step, decision and failure |
| agent registry | Capabilities, data sources, counters, self-ratings, reflections |
| data layer | Live public APIs: PubMed, ClinicalTrials.gov, Open Targets, ChEMBL, Reactome |
| scheduler | Background tasks plus a 15-minute platform tick for schedules and autonomy |
10API reference
All routes are prefixed with /api and return JSON.
| endpoint | description |
|---|---|
| POST /api/runs | Launch a research run {goal, goal_id?} |
| GET /api/runs | List runs (?status, ?goal_id) |
| GET /api/runs/{id} | Run with steps & insights |
| POST /api/runs/{id}/cancel | Cancel a running task |
| GET|POST|PATCH|DELETE /api/goals | Research goals (ACTIVE · PROPOSED · COMPLETED · ARCHIVED) |
| POST /api/goals/{id}/pursue | Launch a run for a goal |
| POST /api/orchestrator/daily-scan | Autonomous daily scan → self-set goal → run |
| POST /api/orchestrator/self-correction | Start a self-correction cycle |
| GET /api/orchestrator/self-corrections | Self-correction history |
| POST /api/debates | Convene a debate {topic, insight_id?} |
| GET /api/debates/{id} | Debate transcript & resolution |
| GET /api/agents | Agent registry |
| POST /api/agents/{key}/reflect | Trigger agent self-reflection |
| GET /api/insights | Insights (?verdict, ?starred) |
| POST /api/insights/{id}/star | Toggle star |
| GET /api/learnings | Strategy memory |
| POST /api/bias/detect | Bias analysis {content, content_type, context?} |
| GET /api/bias/reports/{id} | Bias report |
| POST /api/chat/sessions | New assistant conversation |
| POST /api/chat/sessions/{id}/message | Send message (Server-Sent Events stream) |
| GET|PUT /api/settings | Autonomy, models, alerts |
| GET /api/audit | Audit trail (?agent, ?run_id) |
| GET /api/stats · /api/health | Metrics & system status |
| POST /api/cron/tick | Platform scheduler hook (Bearer secret) |
11FAQ
12Research library
Every completed research run, with its full report: goal, plan, evidence trail, hypotheses, debates, bias audit and lessons learned.