Extracting Metadata from Scanned ETDs

2020

We developed a conditional random field (CRF) model that combines text-based and visual features to automatically extract metadata from scanned Electronic Theses and Dissertations (ETDs). We evaluated our method using 500 human-validated ETD cover pages. The model significantly outperformed text-only and heuristic approaches, achieving 81.3%–96% F1 scores across seven metadata fields.