Documentation

The OpenAlz research foundry, explained.

Everything about how the platform works — from the self-evolving loop and the ten agents to scoring, ethics, the API — plus a live library of every research report the foundry has produced.

01Overview

OpenAlz is a long-horizon, agentic research platform for Alzheimer's disease (AD). A Master Orchestrator coordinates a swarm of specialised sub-agents that continuously read public biomedical data, generate falsifiable hypotheses, verify their own citations, debate one another, audit for bias and — crucially — learn from every cycle so the next one is better.

OpenAlz is built as a single, transparent research service: one orchestrator, nine specialist agents, an ethical reasoning gateway, a bias detection service, an agent registry and an append-only audit trail — all reading five live public databases and writing to one shared evidence store.

10
specialised agents
5
live public data sources
4
self-evolution mechanisms

02Quick start

  1. 1
    Open the foundry

    Go to /app. The Command Center shows live stats and the LLM queue.

  2. 2
    Set or launch a goal

    Type a research goal (e.g. “TREM2 signalling and tau propagation”) and press Launch run — or create a long-horizon goal under Research Goals and press Pursue.

  3. 3
    Watch the agents

    The run page streams every agent step; expand any step to see the raw evidence (PMIDs, NCT IDs, targets, pathways).

  4. 4
    Read the insights

    Each hypothesis shows novelty / confidence / rigor, verified citations, the debate, bias flags and suggested collaborators. Convene further debates from any insight.

  5. 5
    Let it evolve

    Turn on Autonomous mode in Settings for daily scans, and let self-correction run every N runs. Inspect Strategy Memory to see what it learned.

03How self-evolution works

OpenAlz doesn't just answer — it changes itself. Four feedback mechanisms close the loop:

mechanismwhat happenseffect on future runs
run debriefAt the end of each run the orchestrator reviews validator scores, debate resolutions, the bias audit and failed agents, and distils three reusable lessons.The newest 12 lessons are injected into every agent's system prompt.
self-correctionEvery N completed runs (or on demand) the orchestrator reviews agent success/failure counters, recent failures, score distributions and current memory.Writes adaptations to memory, retires stale lessons, proposes new goals.
agent reflectionAny agent can critique its own recent steps: strengths, weaknesses, a self-rating and concrete adjustments.Adjustments are injected into that agent's own instructions.
daily scanIn autonomous mode the orchestrator wakes daily, reads the last 7 days of PubMed and trial updates, and sets or advances a goal.The research agenda follows new evidence without a human prompt.
Scan ─▶ Plan ─▶ Retrieve ─▶ Hypothesise ─▶ Debate ─▶ Audit ─▶ Learn
  ▲                                                          │
  └────────── strategy memory · reflections · goals ◀────────┘

04The agents

agentrolegrounded in
Master OrchestratorDaily scans, goal setting, task decomposition, hypothesis synthesis, debate arbitration, self-correction.LLM reasoning
Literature BridgerSearches AD literature and an orchestrator-chosen adjacent field; finds mechanistic bridges and gaps.PubMed E-utilities
Biomarker HunterProposes fluid / imaging / genetic / digital biomarkers with evidence strength.PubMed + Open Targets
Drug ScreenerShortlists clinical-stage and repurposable drugs; identifies untargeted genes.ChEMBL + Open Targets
Trial OptimizerAnalyses trial landscape, failure patterns, endpoints and cohort gaps.ClinicalTrials.gov v2
Pathway ModelerEnriches the evidence gene set and models pathway crosstalk.Reactome Analysis Service
Data HarmonizerUnifies entities across agents; flags convergent signals & inconsistencies.All agent outputs
Hypothesis ValidatorVerifies citations, raises the strongest objection, scores novelty/confidence/rigor.PubMed + ClinicalTrials.gov
Collaboration MatchmakerMatches insights to real authors from retrieved papers.PubMed author records
Bias Detection ServiceAudits evidence, trial cohorts and hypotheses for bias; powers the Bias Portal.Trial demographics + text

05Data sources

sourceused forendpoint
PubMedLiterature search (5-year window; 7-day for daily scans), abstracts, authors, citation verificationeutils.ncbi.nlm.nih.gov
ClinicalTrials.govAD trial landscape, enrollment, age & sex criteria, NCT verificationclinicaltrials.gov/api/v2
Open TargetsTop AD (MONDO_0004975) target–disease associationsapi.platform.opentargets.org
ChEMBLDrug indications for Alzheimer's diseaseebi.ac.uk/chembl/api
ReactomePathway over-representation for the evidence gene setreactome.org/AnalysisService

06Methodology & scoring

For each run the orchestrator generates three hypotheses grounded only in identifiers present in the retrieved evidence. The Hypothesis Validator then scores each one:

scoremeaningnotes
noveltyHow far the hypothesis is from prior art and from the foundry's own previous insights0–100
confidenceStrength and consistency of the cited evidencePenalised for unverified citations
rigorFalsifiability and quality of the proposed validation experiment0–100
verdictFinal orchestrator ruling after arbitrationaccept · revise · reject

Citation verification: every PMID is checked with PubMed esummary; every NCT ID with the ClinicalTrials.gov API. Each citation is shown with ✓ or ✗ in the UI and report.

07Multi-agent debate

Debates happen inside every run (validator objection → proponent rebuttal → orchestrator resolution) and on demand from the Orchestrator page or any insight. An on-demand debate retrieves fresh PubMed evidence, then runs four turns — proponent (Literature Bridger), skeptic (Hypothesis Validator), rebuttal, counter — before the orchestrator rules supported, contested or refuted with a confidence and next steps. The verdict is attached to the insight's debate history.

08Privacy & ethics

  • No patient data. Only public, aggregate sources are queried.
  • Ethical LLM gateway. Chat prompts are screened for prompt-injection patterns and PII (emails, phone numbers, SSNs are redacted before reaching the model); flags are shown to the user.
  • Full audit trail. Every step start/finish/failure, self-correction and debate resolution is logged.
  • Bias detection. Each run ends with a bias audit; anyone can submit text to the Bias Detection Portal.
  • Limits. LLMs can be wrong; abstracts are truncated; short gene symbols can match unrelated papers; scores are model judgements, not statistics. Not medical advice.

09Architecture

React (landing · /doc · /app)  ──HTTPS──▶  FastAPI /api
                                            ├─ orchestrator + 9 agents (asyncio tasks)
                                            ├─ priority LLM limiter (chat > research)
                                            ├─ autonomy: daily scan · self-correction · debates · reflection
                                            ├─ ethical gateway · bias detection · notifications
                                            └─ MongoDB: runs · insights · goals · learnings · audit_logs …
Platform cron (every 15 min) ──▶ /api/cron/tick ──▶ due schedules + autonomous daily scan
componentresponsibility
orchestratorRun pipeline, daily scan, self-correction, debate arbitration
specialist agents (9)Literature, biomarkers, chemistry, trials, pathways, harmonisation, validation, matchmaking, bias
reasoning gatewaySwappable engines, priority queue (chat before research), prompt-injection & PII screening, streaming
bias detectionBias audit in every run plus the open Bias Detection Portal
audit trailAppend-only log of every step, decision and failure
agent registryCapabilities, data sources, counters, self-ratings, reflections
data layerLive public APIs: PubMed, ClinicalTrials.gov, Open Targets, ChEMBL, Reactome
schedulerBackground tasks plus a 15-minute platform tick for schedules and autonomy

10API reference

All routes are prefixed with /api and return JSON.

endpointdescription
POST /api/runsLaunch a research run {goal, goal_id?}
GET /api/runsList runs (?status, ?goal_id)
GET /api/runs/{id}Run with steps & insights
POST /api/runs/{id}/cancelCancel a running task
GET|POST|PATCH|DELETE /api/goalsResearch goals (ACTIVE · PROPOSED · COMPLETED · ARCHIVED)
POST /api/goals/{id}/pursueLaunch a run for a goal
POST /api/orchestrator/daily-scanAutonomous daily scan → self-set goal → run
POST /api/orchestrator/self-correctionStart a self-correction cycle
GET /api/orchestrator/self-correctionsSelf-correction history
POST /api/debatesConvene a debate {topic, insight_id?}
GET /api/debates/{id}Debate transcript & resolution
GET /api/agentsAgent registry
POST /api/agents/{key}/reflectTrigger agent self-reflection
GET /api/insightsInsights (?verdict, ?starred)
POST /api/insights/{id}/starToggle star
GET /api/learningsStrategy memory
POST /api/bias/detectBias analysis {content, content_type, context?}
GET /api/bias/reports/{id}Bias report
POST /api/chat/sessionsNew assistant conversation
POST /api/chat/sessions/{id}/messageSend message (Server-Sent Events stream)
GET|PUT /api/settingsAutonomy, models, alerts
GET /api/auditAudit trail (?agent, ?run_id)
GET /api/stats · /api/healthMetrics & system status
POST /api/cron/tickPlatform scheduler hook (Bearer secret)

11FAQ

12Research library

Every completed research run, with its full report: goal, plan, evidence trail, hypotheses, debates, bias audit and lessons learned.