Benchmarking Large Language Models Against Tumour Boards for Oncology Treatment Recommendations

⚠️

This is not medical advice. AI-assisted translation — inaccuracies may occur. Always verify the original and consult your oncologist before taking any steps.

About the trial

BEACON (Benchmarking AI for Clinical Oncology decisioNmaking) is a prospective, multicentre, comparative, blinded, non-interventional benchmark evaluating the treatment recommendations of five frontier large language models (LLMs) against the recommendations of multidisciplinary tumour boards (RCP) in oncology treatment planning. One hundred standardised synthetic cases (20 per localisation, across breast, lung, urological, digestive and gynaecological cancers) are submitted as identical structured input to two independent tumour boards per localisation and to five frontier LLMs. Each recommendation - human or model - is decomposed into five predefined decision domains (intent, surgery, radiotherapy, systemic therapy, work-up and biomarkers) and scored 0/1/2 for concordance against a two-tier reference: the consensus of the two tumour boards, complemented by an a priori locked guideline matrix (ESMO, NCCN). The primary endpoint is domain-level concordance between LLM and RCP consensus, expressed as a linearly weighted Cohen's kappa. A co-primary safety endpoint captures the proportion of recommendations carrying serious harm potential, because concordance alone can conceal dangerous errors. Because expert boards may disagree with one another on identical cases, model performance is always interpreted against the human consensus. BEACON is designed as reusable, openly licensed, pre-registered infrastructure: all synthetic cases, evaluation rubrics, the locked guideline matrix, scoring algorithms and verbatim prompts are released for full reproducibility.

Original English text from ClinicalTrials.gov

Who can (and can't) join

✓ Qualifies

  • Dotyczy sprawy raka z jednej z pięciu określonych lokalizacji (pierś, płuca, układ moczowy, układ pokarmowy, ginekologia).
  • Zawiera kompletne dane o stadium choroby, markerach, stanie ogólnym pacjenta i jego chorobach współistniejących.
  • Zawiera pytanie o plan leczenia, na które można odpowiedzieć, kierując się aktualnymi wytycznymi.

✗ Disqualifies

  • Sprawa dotyczy raka innej lokalizacji niż pięć wymienionych.
  • Dane są niekompletne, wewnętrznie sprzeczne lub niejasne.
  • Pytanie o leczenie nie może być rozwiązane zgodnie z obecnymi wytycznymi.

Simplified criteria — AI translation

Trial details

Minimum age
18 Years
Last updated (source)
July 31, 2026
Sex
No restrictions

Therapies / drugs in trial

Locations (1)

Hopital Européen Georges Pompidou

Paris, France

Trial contact

Contact information from ClinicalTrials.gov. Contact in English.

Share this trial

Data from ClinicalTrials.gov. AI-assisted translation, last sync: 8/1/2026.

You can help another person

We fund the Radar, this site, and the development of our Grosz dla Życia fundraising platform from donations and our own resources. Every contribution, even a small one, really helps.