MORGOTH
(2026)Objective
Multicentre development and external validation of MORGOTH, a foundation model for comprehensive automated EEG interpretation across 7 event-level and 17 EEG-level clinical tasks, benchmarked against human experts and state-of-the-art models.
Study Summary
• Event-level detection was especially strong for seizure/ictal–interictal–injury continuum (EUC 96.6%) and spike detection (EUC 100%); on IIIC-Test AUC-ROC 0.990 and on SN2-Test spike detection AUC-ROC/PR 1.00.
• Generalisation drop from internal to external test sets was modest (event-level AUC −1.21%, EUC −3.33%; EEG-level AUC −2.12%, EUC −9.52%) and MORGOTH showed lower age sensitivity (20.90% vs 30.90%) and fewer sex-related differences (33.33% vs 44.00%) than SPaRCNet.
Intervention
MORGOTH — a 19-channel EEG transformer foundation model (tokeniser with 8192 tokens via contrastive learning + vector quantisation; 12 encoder blocks with multihead attention; fully connected event-level heads and convolutional/attention EEG-level heads) — vs board-certified human experts and SOTA models (SPaRCNet, SpikeNet 2, U-Sleep, SCORE-AI, LaBraM, EEGFormer, BrainBERT, Kaggle competition winner).
Inclusion Criteria
EEGs from routine outpatient clinics, epilepsy monitoring units, ICUs, and sleep laboratories; ages 0 to >90 years; expert-labelled event- and EEG-level annotations; development EEGs collected Jan 1, 2003 – Feb 1, 2025 from 4 hospitals (MGH, BWH, BIDMC, Boston Children's Hospital); external validation from 48 institutions across 8 countries.
Study Design
Arms: MORGOTH (AI foundation model) vs human experts (6–30 per test set) vs SOTA baseline models (SPaRCNet, SpikeNet 2, U-Sleep, SCORE-AI, Kaggle winner, LaBraM, EEGFormer, BrainBERT)
Patients per Arm: Development 18 677 patients (14 500 pretraining, 18 677 fine-tuning); internal validation 13 334 patients / 13 618 EEGs; external validation 1573 patients / 1800 EEGs from 48 institutions. Total 12 datasets, 34 602 patients, 52 hospitals, 8 countries.
Outcome
• Experts' operating points under the curve (EUC): outperformed 81.0% (46.9–100; 17/21) of experts on 3 internal multi-expert datasets and 50.8% (14.3–98.0; 18/35) on 4 external multi-expert datasets.
• MoE dataset (21 experts; MoE-Internal 1841 patients/2125 EEGs and MoE-External 409 patients/636 EEGs) — AUC-ROC 0.945 (0.925–0.962), AUC-PR 0.803 (0.731–0.860) across 14 categories; IIIC-Test AUC-ROC 0.990 (0.988–0.991), AUC-PR 0.940 (0.926–0.953), outperforming Kaggle winner by 0.041/0.168 and SPaRCNet by 0.126/0.387; SN2-Test spikes AUC-ROC and AUC-PR both 1.00 (vs SpikeNet 2 0.975/0.990).
• Sleep staging (UPenn): AUC-ROC 0.930 (0.927–0.933) vs U-Sleep 0.917 (0.914–0.921); AUC-PR 0.739 vs 0.706. TELEEEG (23 LMICs, 8099 EEGs): AUC-ROC >0.7 in 9/12 tasks.
• Robustness: lower age sensitivity than SPaRCNet (20.90% vs 30.90%) and fewer sex-related differences than SPaRCNet (33.33% vs 44.00%) and Kaggle winner (48.00%); minimal train–test overfitting (mean 2.75%, never >5%).
• IRR: matched or exceeded expert consensus across 7 multi-expert datasets; on IIIC-Test outperformed experts on all IRR metrics with a substantial lead for LPDs.
Bottom Line
MORGOTH is the first EEG foundation model to deliver expert-level, generalisable performance across the full range of clinically relevant EEG tasks and settings. It matched or exceeded expert consensus and outperformed SOTA task-specific models on internal and external multi-institutional data, with only modest generalisation losses and superior robustness across age, sex, and moderate channel loss, providing a scalable path to expand EEG diagnostic capacity in low-resource and high-volume settings.
Major Points
- Foundation model trained on 14 500 patients (HEEDB pre-training) and fine-tuned on 18 677 patients from 4 Boston-area hospitals (MGH, BWH, BIDMC, BCH), then internally validated on 13 334 patients / 13 618 EEGs and externally validated on 1573 patients / 1800 EEGs from 48 institutions across 8 countries (52 hospitals total, 12 datasets, 34 602 patients).
- Performs 7 event-level tasks (3 binary — normal vs abnormal, burst suppression, spike detection; two 3-class — focal vs generalised vs none for slowing and spike localisation; one 5-class sleep staging; one 6-class seizure/IIIC — seizure, LPDs, GPDs, LRDA, GRDA, other) and 17 EEG-level binary tasks (one per finding).
- Internal validation: average AUC-ROC 0.921 (95% CI 0.780–0.981) across 7 event-level and 17 EEG-level tasks; outperformed on average 81.0% (46.9–100; 17/21) of experts on 3 multi-expert internal datasets.
- External validation: average AUC-ROC 0.908 (0.705–0.980) across 7 event-level and 8 EEG-level tasks; outperformed on average 50.8% (14.3–98.0; 18/35) of experts on 4 multi-expert external datasets; exceeded ≥20% of experts on every one of the 17 evaluated tasks.
- MoE dataset (21 experts; MoE-Internal 1841 patients/2125 EEGs and MoE-External 409 patients/636 EEGs, 14 categories): AUC-ROC 0.945 (0.925–0.962), AUC-PR 0.803 (0.731–0.860), outperforming automated burst-suppression detector by 0.084 and 0.428 (statistically significant). IIIC-Test set (30 experts): outperformed 96.6% (82.1–100; 29/30) of experts, surpassing all in LPD, GPD, LRDA; AUC-ROC 0.990 (0.988–0.991), AUC-PR 0.940 (0.926–0.953), exceeding Kaggle winner by 0.041/0.168 and SPaRCNet by 0.126/0.387 (per Figure 2B per-class averages).
- SN2-Test (24 experts): outperformed all 24 experts in spike detection; MORGOTH AUC-ROC and AUC-PR both 1.00 vs SpikeNet 2 (0.975 / 0.990).
- External generalisation drop was small — event-level AUC −1.21% (0.00–3.94), EUC −3.33% (0.00–33.33); EEG-level AUC −2.12% (0.00–5.83), EUC −9.52% (0.00–75.00). Overfitting minimal: mean train–test accuracy difference 2.75%, never >5%.
- Sleep staging: on external UPenn (6 experts) AUC-ROC 0.930 (0.927–0.933) vs U-Sleep 0.917 (0.914–0.921); AUC-PR 0.739 vs 0.706, despite U-Sleep using an added EOG channel. On MASS dataset outperformed U-Sleep by average AUC-ROC 0.030 (0.022–0.038). N1 remained hardest (AUC-PR 0.262).
- Reliability (Cohen's κ): matched or exceeded expert consensus on 7 multi-expert datasets. On IIIC-Test outperformed experts on all IRR metrics with substantial LPD lead. On SN2-Test achieved moderate agreement with individual experts and near-perfect agreement with consensus. On TELEEEG (23 LMICs, 8099 EEGs) AUC-ROC >0.7 in 9 of 12 tasks.
- Robustness across demographics: on IIIC-Test lower age sensitivity (20.90% vs 30.90%) and fewer sex-related differences (33.33% vs 44.00% SPaRCNet; 48.00% Kaggle winner). Local features more affected by missing channels than global patterns.
Study Design
- Study Type
- Multicentre retrospective development and external validation of an AI foundation model for comprehensive automated EEG interpretation (diagnostic accuracy study; not a randomised interventional trial)
- Randomization
- No
- Blinding
- Expert annotators blinded to model outputs; multi-expert labels obtained independently and used as consensus ground truth. IRR analysis compared expert–expert with expert–model pairs.
- Sample Size
- 34602
- Follow-up
- Not applicable (diagnostic study; no longitudinal patient follow-up)
- Centers
- 52
- Countries
- USA, Canada, Belgium, Switzerland, Brazil, UK, China, Greece
Primary Outcome
Definition: Overall diagnostic performance across 7 event-level and 17 EEG-level tasks, measured by AUC-ROC, AUC-PR, and experts' operating points under the curve (EUC), on internal (HEEDB-Test, IIIC-Test, SN2-Test, MGH-PSG-Test, BCH-PSG, MoE-Internal) and external (HEP, SAI, ON, TUH-Test, UPenn, MASS, MoE-External, TELEEEG) validation datasets
| Control | Intervention | HR/OR | P-value |
|---|---|---|---|
| - | - | - | - |
Limitations & Criticisms
- Uses only the standard 19 EEG channels; does not incorporate EMG, EOG, or other polysomnography signals — limits sleep-related task performance and detection of arousals, apnoeas, and limb movements.
- Training data drawn mainly from tertiary US epilepsy centres (MGH, BWH, BIDMC, BCH) — may reflect referral-centre case-mix biases.
- Does not explicitly detect certain normal variants or artifacts (breach rhythm, photoparoxysmal responses) that even experts interpret variably.
- Model distinguishes focal vs generalised activity but does not localise focal activity or distinguish between seizure subtypes — limits presurgical utility.
- Does not use clinical context (e.g., treatment response) — focuses on pattern recognition rather than outcome prediction (epilepsy risk or syndrome classification).
- External validation drop for EEG-level tasks larger than for event-level (EUC −9.52% vs −3.33%), suggesting some site-specific overfitting for whole-recording classifications.
- N1 sleep stage detection remained hard (AUC-PR 0.262; no improvement post-calibration), reflecting expert-level ambiguity.
- Human–AI workflow (preliminary reports generated for clinician review) not prospectively validated — prospective trials still needed for regulatory approval.
- TELEEEG (LMIC) performance declined (some tasks AUC-ROC <0.7) due to missing channels and label noise, indicating need for further expert review in resource-limited real deployment.
- Study is retrospective diagnostic accuracy work — not a randomised interventional trial; does not directly measure impact on patient care or outcomes.
Citation
Lancet Digit Health 2026; published online. DOI: 10.1016/j.landig.2026.101039