ToothFairy4

Multicenter CBCT dataset with 625 cases and bilingual clinical reports for AI dental and surgical-planning documentation.

Overview

This page describes the ToothFairy4 dataset, the latest extension of the ToothFairy series developed at the University of Modena and Reggio Emilia with the collaboration of Radboud University and the Digital Research Center of Sfax. The dataset is part of the ODIN challenges included in the ODIN2026 workshop (MICCAI2026). For challenge details, visit the dedicated website or check the structured submission document.


Dataset Focus

Clinical Reasoning: reports are designed to train systems capable of identifying dental status, bone quality and quantity, and anatomical risks to generate structured reports for procedures such as implant placement, extractions, and sinus lifts.

Multimodal Integration: the core technical objective is to convert high-dimensional 3D medical imaging (CBCTs) into coherent clinical documentation.

Generalization: the benchmark is designed to assess robustness across centers and acquisition protocols, with a hidden private test set from an external institution.

 

Dataset Structure

Each case is organized as:

  • cbct/: one CBCT volume (NIfTI, compressed as .nii.gz),
  • reports_it/: original clinician-authored report(s) in Italian,
  • reports_en/: English-translated report(s).

The dataset is divided into four subsets: P, F, S, and A. Sets P, F, and S come from ToothFairy3 (be careful about the volume orientation if you want to reuse the segmentation labels from ToothFairy3), while set A is the new subset contributed by the Digital Research Center of Sfax, acquired with different CBCT machines. While other sets ensure a 0.3 mm isotropic pixel spacing, set A does not. The number of reports available per patient ranges from a single report (250 patients) to 3 (for only 3 patients). The majority of cases (371 patients) have exactly two reports redacted by different maxillofacial surgeons with >5 years of experience.

Field Value
Total CBCT/Patients 625 patients
Set P 417 volumes
Set F 63 volumes
Set S 52 volumes
Set A 100 volumes
CBCT files per case 1 volume (.nii.gz)
Reports (IT and EN) 1001 total for each language
Reports per case (IT and EN) 250 cases with 1 report, 371 with 2 reports, 3 with 3 reports

Clinical Context

ToothFairy4 is designed around real reporting workflows in implantology and oral-maxillofacial surgery, where clinicians must translate complex CBCT findings into clear, structured planning notes. In these settings, reports often need to summarize dentition status, bone characteristics, anatomical constraints, and risk-related structures in a concise but clinically meaningful way.

The benchmark focuses on generating clinically useful draft reports from CBCT data, with an emphasis on factual completeness, anatomical consistency, and reliable performance across centers. By evaluating how well models capture the information needed for implant and surgical planning, ToothFairy4 aims to support more standardized clinical documentation workflows.

Translation Pipeline

Original reports (in Italian) are preserved, and English versions are produced with an LLM-based translation pipeline. The current translation script uses gpt-5.2.

To mitigate translation errors, random subsets of training reports are quality checked by clinicians. 

The test-set translation, i.e., those used for the ToothFairy4 challenge ranking and not publicly released, is clinically reviewed to ensure no clinical errors occurred in the LLM-based translation.


System Prompt 

The prompt used for report translation is the following:

You are an expert bilingual medical translator specializing in dental radiology and CBCT reporting.

Task: translate an Italian clinical dental/CBCT report into English.

Rules:
1. Preserve meaning with maximal clinical fidelity. Do not add, remove, soften, or overstate findings.
2. Use natural, native-level medical English as written by a radiologist fluent in Italian.
3. Keep the narrative form and sentence flow of the source report (no bullet points, no restructuring into templates).
4. Preserve all clinical qualifiers and uncertainty language faithfully (e.g., possible, probable, compatible with).
5. Preserve tooth numbering, side labels, measurements, anatomic references, and abbreviations exactly when clinically appropriate.
6. Keep register professional and concise.
7. Return only the final English translation text, with no preface or commentary.

Download

You need to have an account to download the dataset. Please sign in or sign up!

How to cite

If you use our dataset, you must cite the following papers.

text loading...

Copy