Johnson, Jon and De, Suparna and Oldroyd, Becky and White, Sarah and Mills, Hayley and Pravin, Chandresh (2026). Extraction and Utilisation of Metadata from Non-machine-actionable Documents to Improve Data Curation and Discovery, 2024-2025. [Data Collection]. Colchester, Essex: UK Data Service. 10.5255/UKDA-SN-855570
The proposal builds on prior exploratory work on the use of machine learning to automate data and metadata curation for survey data, to help overcome the current reliance on non-sustainable manual processes. Such automation will also provide additional metadata which can be used to improve the discovery, evaluation and curation of these rich and widely used research investments, and be the basis for further innovations such as automation of disclosure risk control to support the needs of an emerging FAIR research landscape.
The proposal uses CLOSER Discovery as a training dataset to develop machine learning (ML) models for the extraction of metadata from survey questionnaires, and its annotation to an established vocabulary to support discovery. It will use the Growing Up in Scotland study as its metadata and data sources. The project will also develop ML models for the identification of key variables for input into the Anonymisation Decision Making Framework and automation of disclosure risk assessment.
To achieve these aims, the proposal outlines our approach to the development of novel machine learning models which are tailored to the specific challenges of semantically rich survey data collection and research datasets. The proposal outlines how the alignment of both structural (standards) and semantic metadata (controlled vocabularies and conceptual frameworks) as the output from these ML models can be used to create metadata resources which meet the evolving needs of researchers from a range of disciplines who utilise longitudinal population survey (LPS) and other survey data.
The proposal will bring immediate benefits to the interoperability of CLOSER Discovery with the UKDS, and the Consortium of European Social Science Data Archives (CESSDA); provide a prototype for further development in extracting high quality metadata from the UKDS archive of survey questionnaires in PDF; provide a prototype for creating metadata resources to provide harmonisable data for LPS and create robust ML models for supporting metadata curation pipelines and more robust and scalable disclosure risk assessment of survey data.
The proposal brings together a team with expertise in survey data collection, metadata production and curation, computer science, data archiving and disclosure risk assessment to develop methods which address the key challenges in the call to support better data discovery, provide the basis for more automated, accurate, and faster approaches to curating data.
Data description (abstract)
This collection builds upon exploratory work into the use of machine learning to automate data and metadata curation. The primary motivation for this work is to overcome the current reliance on non-sustainable, manual curation processes. By developing automated methods, the project aims to improve the discovery, evaluation, and curation of longitudinal population survey data, whilst establishing a foundation for innovations such as automated disclosure risk control within a FAIR data landscape.
The origin PDF, a paper version of the CAPI questionnaire from the National Child Development Study Biomedical Sweep in 2002-2004 is available in the data collection "University of London, Institute of Education, Centre for Longitudinal Studies, NatCen Social Research. (2024). National Child Development Study: Biomedical Survey 2002-2004. [data collection]. UK Data Service. SN: 8731, DOI: http://doi.org/10.5255/UKDA-SN-8731-1", see related resources.
The deposit comprises of the following:
The file JSON holds the output from a machine learning model processing a PDF into a structured JSON file. The output describes the location of the bounding box for the text, its type and the accompanying extracted text,
The CSV file holds a subset of the information in the JSON file and a confidence estimate for each extracted item, 1.0 being the highest level of confidence.
Based on this work, the CLOSER Topic Vocabulary was enhanced and reorganised, see related resources.
| Data creators: |
|
|||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Contributors: |
|
|||||||||||||||||||||
| Sponsors: | Economic and Social Research Council | |||||||||||||||||||||
| Grant reference: | ES/Z502935/1 | |||||||||||||||||||||
| Topic classification: |
Science and technology Demography (population, vital statistics and censuses) |
|||||||||||||||||||||
| Keywords: | AUTOMATION, ARTIFICIAL INTELLIGENCE, INFORMATION RETRIEVAL | |||||||||||||||||||||
| Grant holders: | Jon Johnson, Deirdre Lungley, Suparna De | |||||||||||||||||||||
| Project dates: |
|
|||||||||||||||||||||
| Date published: | 01 Sep 2026 13:20 | |||||||||||||||||||||
| Last modified: | 01 Sep 2026 13:20 | |||||||||||||||||||||
Downloads
Altmetric
Related Resources
Data collections
Publications
| METACURATE-ML: Generalising the extraction of questionnaires to DDI-Lifecycle. |
Software
| CLOSER Vocabulary (Version 1.0) |
Website
| Extraction and Utilisation of Metadata from Non-machine-actionable Documents to Improve Data Curation and Discovery |
| Publications from project |

