Orthodontic dataset of 1,000 patient cases pairing 3D intra-oral scans and 2D photographs with clinician-authored reports.
The Bite2Text dataset is a comprehensive orthodontic resource comprising 1,000 patient cases. Each case includes 3D intra-oral scans and 2D photographs, paired with clinician-authored reports on malocclusion and treatment findings. This multi-center dataset includes a hidden test set, making it ideal for developing and evaluating multimodal models that generate clinically meaningful orthodontic reports. The paper describing the dataset has been published at ECCV2026 and is available here. This benchmark is developed by the University of Modena and Reggio Emilia and the University of Ferrara.

Orthodontic Reporting. The dataset allows the development of models able to generate clinically meaningful reports describing malocclusion patterns, occlusal relationships, crowding or spacing, and treatment-relevant findings.
Multimodal Reasoning. 3D intraoral geometry and 2D intraoral photographs should be jointly processed to produce coherent, clinically aligned text.





Data is structured to facilitate easy access and analysis for research and clinical applications. Each case/patient is organized as:
ios/: upper and lower 3D intraoral scans,photos/: corresponding standardized intraoral photographs,reports_it/: original clinician-authored reports,reports_en/: English-translated reports.Relevant dataset details are reported in the table to the right.
| Field | Value |
| Cases | 1,000 patients |
| 3D Input | Upper and lower intraoral scans (IOS) |
| 2D Input | Standardized intraoral RGB photographs |
| Annotation format | Clinician-authored textual reports (original and English translation) |
Bite2Text reflects routine orthodontic workflows where clinicians inspect intraoral scans and photographs to describe occlusal relationships, alignment patterns, and treatment-relevant anomalies.
The objective is to support high-quality draft reporting that reduces manual documentation time, improves inter-observer consistency, and enables reliable multi-center clinical decision support.
The training set aggregates acquisitions from multiple centers, scanners, and clinical environments.
Ethical Approvals: data release is covered by local approvals from the University of Ferrara (No. 262/2025/Oss/UniFe).
If you use our dataset, you must cite the following papers.
text loading...
Copy