← Back
NeuroTrials.ai
Neurology Clinical Trial Database

MORGOTH

Toward unified and comprehensive automated electroencephalogram interpretation: a multicentre development and validation of an electroencephalogram foundation model

Year of Publication: 2026

Authors: Sun C, Karakis I, Herlopian A, ..., Jing J)

Journal: Lancet Digital Health

Citation: Lancet Digit Health 2026; published online. DOI: 10.1016/j.landig.2026.101039

Link: https://doi.org/10.1016/j.landig.2026.101039


Clinical Question

Can a single AI foundation model deliver expert-level, comprehensive automated EEG interpretation across all major clinical settings (routine outpatient, epilepsy monitoring units, ICU, sleep labs) and across the full spectrum of clinically relevant patterns (seizures, epileptiform discharges, rhythmic and periodic patterns of the ictal–interictal–injury continuum, pathological slowing, sleep staging)?

Bottom Line

MORGOTH is the first EEG foundation model to deliver expert-level, generalisable performance across the full range of clinically relevant EEG tasks and settings. It matched or exceeded expert consensus and outperformed SOTA task-specific models on internal and external multi-institutional data, with only modest generalisation losses and superior robustness across age, sex, and moderate channel loss, providing a scalable path to expand EEG diagnostic capacity in low-resource and high-volume settings.

Major Points

  • Foundation model trained on 14 500 patients (HEEDB pre-training) and fine-tuned on 18 677 patients from 4 Boston-area hospitals (MGH, BWH, BIDMC, BCH), then internally validated on 13 334 patients / 13 618 EEGs and externally validated on 1573 patients / 1800 EEGs from 48 institutions across 8 countries (52 hospitals total, 12 datasets, 34 602 patients).
  • Performs 7 event-level tasks (3 binary — normal vs abnormal, burst suppression, spike detection; two 3-class — focal vs generalised vs none for slowing and spike localisation; one 5-class sleep staging; one 6-class seizure/IIIC — seizure, LPDs, GPDs, LRDA, GRDA, other) and 17 EEG-level binary tasks (one per finding).
  • Internal validation: average AUC-ROC 0.921 (95% CI 0.780–0.981) across 7 event-level and 17 EEG-level tasks; outperformed on average 81.0% (46.9–100; 17/21) of experts on 3 multi-expert internal datasets.
  • External validation: average AUC-ROC 0.908 (0.705–0.980) across 7 event-level and 8 EEG-level tasks; outperformed on average 50.8% (14.3–98.0; 18/35) of experts on 4 multi-expert external datasets; exceeded ≥20% of experts on every one of the 17 evaluated tasks.
  • MoE dataset (21 experts; MoE-Internal 1841 patients/2125 EEGs and MoE-External 409 patients/636 EEGs, 14 categories): AUC-ROC 0.945 (0.925–0.962), AUC-PR 0.803 (0.731–0.860), outperforming automated burst-suppression detector by 0.084 and 0.428 (statistically significant). IIIC-Test set (30 experts): outperformed 96.6% (82.1–100; 29/30) of experts, surpassing all in LPD, GPD, LRDA; AUC-ROC 0.990 (0.988–0.991), AUC-PR 0.940 (0.926–0.953), exceeding Kaggle winner by 0.041/0.168 and SPaRCNet by 0.126/0.387 (per Figure 2B per-class averages).
  • SN2-Test (24 experts): outperformed all 24 experts in spike detection; MORGOTH AUC-ROC and AUC-PR both 1.00 vs SpikeNet 2 (0.975 / 0.990).
  • External generalisation drop was small — event-level AUC −1.21% (0.00–3.94), EUC −3.33% (0.00–33.33); EEG-level AUC −2.12% (0.00–5.83), EUC −9.52% (0.00–75.00). Overfitting minimal: mean train–test accuracy difference 2.75%, never >5%.
  • Sleep staging: on external UPenn (6 experts) AUC-ROC 0.930 (0.927–0.933) vs U-Sleep 0.917 (0.914–0.921); AUC-PR 0.739 vs 0.706, despite U-Sleep using an added EOG channel. On MASS dataset outperformed U-Sleep by average AUC-ROC 0.030 (0.022–0.038). N1 remained hardest (AUC-PR 0.262).
  • Reliability (Cohen's κ): matched or exceeded expert consensus on 7 multi-expert datasets. On IIIC-Test outperformed experts on all IRR metrics with substantial LPD lead. On SN2-Test achieved moderate agreement with individual experts and near-perfect agreement with consensus. On TELEEEG (23 LMICs, 8099 EEGs) AUC-ROC >0.7 in 9 of 12 tasks.
  • Robustness across demographics: on IIIC-Test lower age sensitivity (20.90% vs 30.90%) and fewer sex-related differences (33.33% vs 44.00% SPaRCNet; 48.00% Kaggle winner). Local features more affected by missing channels than global patterns.

Design

Study Type: Multicentre retrospective development and external validation of an AI foundation model for comprehensive automated EEG interpretation (diagnostic accuracy study; not a randomised interventional trial)

Randomization:

Blinding: Expert annotators blinded to model outputs; multi-expert labels obtained independently and used as consensus ground truth. IRR analysis compared expert–expert with expert–model pairs.

Enrollment Period: Jan 1, 2003 – Feb 1, 2025 (EEG data collection window)

Follow-up Duration: Not applicable (diagnostic study; no longitudinal patient follow-up)

Centers: 52

Countries: USA, Canada, Belgium, Switzerland, Brazil, UK, China, Greece

Sample Size: 34602

Power Calculation: Not reported (diagnostic model validation; 95% CIs derived by 10 000 bootstrap iterations; differences considered significant when intervals did not overlap)

Analysis: Performance evaluated with ROC and precision-recall curves. Expert benchmarking used experts' operating points under the curve (EUC) — fraction of expert operating points below the model's ROC/PR curve. Cohen's κ for reliability. Calibration indices derived by parametric-model fitting (range −1 undercalibrated to +1 overcalibrated). Platt scaling and isotonic regression for post-hoc calibration. Training: focal loss with curriculum learning; AdamW optimiser with learning-rate scheduling and early stopping. Two-stage training — self-supervised pretraining on 14 500 HEEDB patients, then task-specific fine-tuning heads. Bootstrap 10 000 iterations for 95% CIs.


Inclusion Criteria

  • Clinical EEG recordings from routine outpatient clinics, epilepsy monitoring units, intensive care units, or sleep laboratories
  • All ages (0 to >90 years) accepted
  • Expert-labelled event-level and/or EEG-level annotations available
  • Recordings compatible with 19-channel 10–20 system input; other sampling rates resampled to 200 Hz; missing channels imputed by mean of available channels
  • Institutional review board approval at contributing sites; waiver of informed consent for retrospective analysis (BIDMC IRB protocols #2022P000481 and #2022P000417)

Exclusion Criteria

  • No explicit patient or EEG exclusion criteria reported by the authors; the study retrospectively included the available expert-annotated recordings from participating institutions

Baseline Characteristics

CharacteristicDevelopment / Fine-tuning (n=18 677 patients)Internal validation (n=13 334 patients / 13 618 EEGs)External validation (n=1573 patients / 1800 EEGs)
Median age (range)56 (0 to >90)59 (0 to >90)45 (0 to >90)
Female n (%)49%49%49%
Male n (%)51%51%51%
Hospitals4 (MGH, BWH, BIDMC, BCH)
10-s labelled event-level segments102 442
10-min EEGs for training 17 EEG-level heads18 677
Ethnicity63.51% White, 7.92% Black, 0.14% American Indian/Alaska Native, 2.88% Asian, 0.04% Native Hawaiian/Pacific Islander, 0.91% Multiracial, 5.65% Other
Setting mix5666 routine, 5340 ICU, 612 EMU, 2269 sleep laboratory, 1020 not available
Institutions48
DatasetsHEP (143), SAI (100), ON (100), TUH-Test (552), UPenn (69), MASS (200), MoE-External (409 patients / 636 EEGs); plus TELEEEG (8099 EEGs, 23 LMICs)

Arms

FieldMORGOTH (AI foundation model)ControlControl
Intervention19-channel EEG transformer foundation model: tokeniser (8192 tokens via contrastive learning + vector quantisation) + 12 encoder blocks with multihead attention and temporal/spatial positional encoding + task heads. Event-level heads use fully connected layers on 10-s or 1-s segments; EEG-level heads use convolutional/attention on 10-min or full recordings. Supports variable-length input at inference.Board-certified neurologists, epileptologists, and sleep medicine specialists — 6 to 30 per test set (MoE 21; IIIC 30; SN2 24; SAI 14; ON 15; UPenn 6; HEP 14). Multi-expert consensus used as reference; individual experts also compared via EUC and Cohen's κ.SpikeNet 2 (spike detection); SPaRCNet and Kaggle competition winner (seizure/IIIC classification); U-Sleep (sleep staging); an automated burst-suppression detector; SCORE-AI (EEG-level tasks); and EEG foundation models LaBraM, EEGFormer, BrainBERT on TUH benchmarks.
N

Outcomes

OutcomeTypeControlInterventionHR / OR / RRP-value
Overall diagnostic performance across 7 event-level and 17 EEG-level tasks, measured by AUC-ROC, AUC-PR, and experts' operating points under the curve (EUC), on internal (HEEDB-Test, IIIC-Test, SN2-Test, MGH-PSG-Test, BCH-PSG, MoE-Internal) and external (HEP, SAI, ON, TUH-Test, UPenn, MASS, MoE-External, TELEEEG) validation datasetsPrimaryMORGOTH internal validation (average across 7 event-level + 17 EEG-level tasks): AUC-ROC 0.921 · MORGOTH external validation (average across 7 event-level + 8 EEG-level tasks): AUC-ROC 0.908 · AUC-ROC 95% CI (internal): 0.780–0.981 · AUC-ROC 95% CI (external): 0.705–0.980
IIIC-Test (30 experts) AUC-ROC — average across 6 classes (per Figure 2B)SecondaryMORGOTH: 0.990 (95% CI 0.988–0.991) · Kaggle winner: 0.949 · SPaRCNet: 0.864 · Notes: MORGOTH exceeded Kaggle winner by 0.041 and SPaRCNet by 0.126 (Figure 2B per-class averages)
IIIC-Test AUC-PR — average across 6 classesSecondaryMORGOTH: 0.940 (0.926–0.953) · Kaggle winner: 0.772 · SPaRCNet: 0.553 · Notes: MORGOTH exceeded Kaggle winner by 0.168 and SPaRCNet by 0.387
IIIC-Test EUC — % of experts outperformedSecondaryMORGOTH: 96.6% (82.1–100; 29/30) · Notes: surpassed all experts in LPD, GPD, LRDA
MoE dataset (21 experts, 14 categories; MoE-Internal 1841 patients/2125 EEGs and MoE-External 409 patients/636 EEGs) AUC-ROCSecondaryMORGOTH: 0.945 (0.925–0.962) · Automated burst suppression detector delta: +0.084 (significant)
MoE AUC-PRSecondaryMORGOTH: 0.803 (0.731–0.860) · Delta vs burst-suppression detector: +0.428 (significant)
MoE experts outperformedSecondaryMORGOTH: 81.0% (46.9–100; 17/21) on 13 of 14 tasks
SN2-Test spike detection (24 experts) AUC-ROC / AUC-PRSecondaryMORGOTH: 1.00 / 1.00 · SpikeNet 2: 0.975 / 0.990 · Notes: outperformed all 24 experts (SpikeNet 2 values from Figure 2C)
Sleep staging — UPenn dataset (6 experts) AUC-ROCSecondaryMORGOTH: 0.930 (0.927–0.933) · U-Sleep: 0.917 (0.914–0.921)
Sleep staging — UPenn AUC-PRSecondaryMORGOTH: 0.739 (0.730–0.748) · U-Sleep: 0.706 (0.698–0.715) · Notes: U-Sleep additionally used an EOG channel; N1 AUC-PR 0.262 (0.248–0.278) — MORGOTH still outperformed 52.7% of experts; REM AUC-PR 0.857 (0.849–0.864) — did not surpass any
MASS sleep dataset — average AUC-ROC advantage over U-SleepSecondaryMORGOTH: +0.030 (0.022–0.038)
External single-scored corpora — AUC-ROCSecondaryTUH Seizure Corpus: 0.883 (0.881–0.885) · TUH EEG Events Corpus: 0.729 (0.695–0.763) · TUH Slowing Corpus: 0.747 (0.715–0.779) · TUH Abnormal EEG Corpus: 0.902 (0.900–0.904)
SAI dataset (14 experts, 5 tasks) — average EUC across tasksSecondaryMORGOTH ROC EUC: 50.62% (24.16–77.14) · MORGOTH PR EUC: 53.50% (26.36–77.58) · Notes: outperformed half of experts
ON dataset (15 experts) — EUC-ROC / EUC-PRSecondaryMORGOTH ROC EUC: 43.8% (14.3–78.6) · MORGOTH PR EUC: 39.0% (8.9–75.0)
HEP dataset — event-level AUC-ROC / AUC-PRSecondaryMORGOTH: 0.999 / 0.999 · Notes: near-perfect; surpassed SpikeNet 2 at both event and EEG level
Generalisation drop (internal → external)SecondaryEvent-level AUC drop: −1.21% (0.00–3.94) · Event-level EUC drop: −3.33% (0.00–33.33) · EEG-level AUC drop: −2.12% (0.00–5.83) · EEG-level EUC drop: −9.52% (0.00–75.00)
IRR on IIIC-Test (Cohen's κ)SecondaryMORGOTH: outperformed experts on all IRR metrics with substantial lead for LPDs
IRR on SN2-TestSecondaryMORGOTH: moderate agreement with individual experts; near-perfect agreement with consensus
TELEEEG (8099 EEGs from 23 LMICs)SecondaryMORGOTH: AUC-ROC >0.7 in 9 of 12 tasks; some decline attributed to missing channels and label noise
Not applicable — diagnostic model, no patient interventions or adverse events reported.SafetyNotes: Overfitting monitored as internal-to-external accuracy delta: mean 2.75%, never exceeded 5%.

Subgroup Analysis

Prespecified subgroup analyses across age and sex on IIIC-Test showed MORGOTH's demographic robustness exceeded baselines: age sensitivity 20.90% vs 30.90% for SPaRCNet; sex-related differences 33.33% (MORGOTH) vs 44.00% (SPaRCNet) and 48.00% (Kaggle winner). Local features (e.g., spike detection) most stable across age; slowing detection showed greater variability. Local features more affected by missing channels than global patterns.


Criticisms

  • Uses only the standard 19 EEG channels; does not incorporate EMG, EOG, or other polysomnography signals — limits sleep-related task performance and detection of arousals, apnoeas, and limb movements.
  • Training data drawn mainly from tertiary US epilepsy centres (MGH, BWH, BIDMC, BCH) — may reflect referral-centre case-mix biases.
  • Does not explicitly detect certain normal variants or artifacts (breach rhythm, photoparoxysmal responses) that even experts interpret variably.
  • Model distinguishes focal vs generalised activity but does not localise focal activity or distinguish between seizure subtypes — limits presurgical utility.
  • Does not use clinical context (e.g., treatment response) — focuses on pattern recognition rather than outcome prediction (epilepsy risk or syndrome classification).
  • External validation drop for EEG-level tasks larger than for event-level (EUC −9.52% vs −3.33%), suggesting some site-specific overfitting for whole-recording classifications.
  • N1 sleep stage detection remained hard (AUC-PR 0.262; no improvement post-calibration), reflecting expert-level ambiguity.
  • Human–AI workflow (preliminary reports generated for clinician review) not prospectively validated — prospective trials still needed for regulatory approval.
  • TELEEEG (LMIC) performance declined (some tasks AUC-ROC <0.7) due to missing channels and label noise, indicating need for further expert review in resource-limited real deployment.
  • Study is retrospective diagnostic accuracy work — not a randomised interventional trial; does not directly measure impact on patient care or outcomes.

Funding

US National Institutes of Health (multiple grants including RFG064312, RF1NS120947, R01AG073410, R01HL161253, R01NS126282, R01AG073598, R01NS117904, R01NS131347, R01NS130119, R01NS111022, R01NS131967, K23AG063899, NS121559). Additional support from: Swebilius Foundation; CONDA Award; DHHS LB606 Nebraska Stem Cell Grant; Fonds National pour la Recherche Scientifique; Brussels Institute for Research and Innovation (INNOVIRIS); Fonds Erasme pour la Recherche Médicale; Fonds Jaumotte; Department of Defense (HT9425-23-1-0242, HT9425-25-1-0170, W81XWH-19-1-0861, W81XWH-21-C-0075); AMFDP 843457 (20CDA35310297 and 24DIVSUP1274116); Regents of the University of California; Cures Within Reach (2022CAL-Amorim); Zoll Foundation; Hellman Foundation; NINDS (K23NS124656, K23NS112596, R01NS117904, SR21NS137117); Brain Aneurysm Foundation; and Veterans Affairs Office of Research and Development (I01HX003107-01A2).

Based on: MORGOTH (Lancet Digital Health, 2026)

Authors: Sun C, Karakis I, Herlopian A, ..., Jing J)

Citation: Lancet Digit Health 2026; published online. DOI: 10.1016/j.landig.2026.101039

Content summarized and formatted by NeuroTrials.ai.