Identifying research data
The OECD defines research data as factual records, including numbers, text, images and sounds, used as primary sources for scientific research and generally recognised by the scientific community as necessary to validate research findings (OECD, 2007, p. 13).
In other words, research data are the materials collected, observed, produced or reused to answer a research question, support an interpretation or verify research findings. Almost all disciplines and fields of research produce data, from mathematics and anthropology to computer science, the humanities and law.
Research data can take many different forms and can be produced using a wide range of methods and instruments.
- Documents (paper or digital), spreadsheets
- Photographs, images, films, video or audio recordings
- Survey responses, transcripts, correspondence tables
- Artefacts, samples
- Laboratory notes, field notes, logbooks
- Computer code, algorithms, models, scripts
- Bibliographies, text or archival corpora
More than its format, the role that a piece of information plays within a research project is a key factor in determining whether it constitutes research data. For example, a photograph, bibliography, text or spreadsheet may be considered research data when it is analysed or directly supports the research findings. In another context, the same materials might simply be administrative documents or communication supports.
In addition, some information may not constitute research data in the strict sense, but is nevertheless essential for understanding, verifying or reusing the data. This includes metadata, README files, protocols, instrument settings, and similar documentation
Data classifications
Several complementary classifications can be used to help identify and characterise the different types of data involved in a research project.
Research data and materials can exist in either physical or digital form. Digital data can also be distinguished according to whether they were created digitally from the outset or produced by digitising a physical original.
- Physical data are data, documents or materials preserved in a tangible, physical form. They can be consulted without the use of a computer or digital environment.
Examples: manuscripts, field notebooks, paper questionnaires, works of art, biological or geological samples, specimens, etc.
- Digital data are data recorded in a form that can be read and processed using a digital environment. They may be created directly in digital form or result from the digitisation of a physical document or object.
Examples: online questionnaire results, databases, spreadsheets, digital photographs, files generated by laboratory instruments, scans of archival documents, digital reproductions of works of art, etc.
The University of Bristol distinguishes five categories of research data according to how they are produced and the extent to which they can be reproduced:
- Observational data are collected in real time through the observation or measurement of phenomena in a specific context.Because the conditions in which they are collected may be difficult or impossible to recreate, these data are often unique and cannot be replaced.
Examples: neuroimaging data, survey results, field measurements and photographs.
- Experimental data are generated through experiments using laboratory instruments, equipment or standardised procedures. They can generally be reproduced, although doing so may require substantial time, effort and resources..
Examples: genomic sequencing, chromatograms, animal experiments
- Simulated or modelled data are generated using computational, mathematical or experimental models. In many cases, the models and the methods used to produce the data are more important for reproducibility than the resulting datasets themselves.
Examples: outputs from economic or climate models
- Derived or compiled data are created by transforming, processing, combining or aggregating existing data.
Examples: datasets produced through text or data mining, compiled databases
- Reference data are curated datasets, collections or corpora that are maintained and used as recognised reference resources within a particular field.
Examples: gene databases, archival collections, historical image databases
Data can be classified according to whether they were generated specifically for the current project or already existed before the project began.
- Primary data are data collected or generated as part of the current project to address its research questions
- Secondary data are pre-existing data reused for the current project. They were originally collected or generated in another context, either by third parties or as part of an earlier project by the same researcher or research team.
Examples: data obtained from a repository or data from a previous research project
- Source or raw data: data preserved in a form that is as close as possible to their original state at the time of collection or generation. In some cases, these data cannot be analysed directly and must first be prepared or processed. The term raw data should be used with caution as these data may already have been converted, filtered or otherwise processed by an instrument or software.
Examples: audio or video recordings, questionnaire responses, photographs, field notes, survey data.
- Prepared or processed data: source data that have been transformed to make them suitable for analysis. This may include cleaning, checking, validating, calibrating, transcribing, annotating, pseudonymising, anonymising, combining or aggregating data.
Examples: verified interview transcripts, cleaned and anonymised questionnaire responses, datasets with harmonised formats or units, calibrated or cropped images.
- Analysis data: data generated through the analysis of source or prepared data.
Examples: coded corpora, analysis matrices, calculated variables, classifications, intermediate results, model outputs, statistical tables.
How to create an inventory of your research data?
The classifications above can help you identify the different types of data involved in your project. The following questions can also help you build a comprehensive overview of your research data:
- What files, documents, observations, objects or materials have been collected or generated?
- What data or materials underpin your analyses, interpretations and conclusions?
- What existing data are being reused from partners, archives, repositories, databases or other sources?
- What data or materials would be needed to verify or reproduce your results?
- What new versions or datasets have been created through cleaning, transcription, annotation, combination, aggregation or analysis?
- What documentation is needed to understand the data and the processing steps applied to them?
Creating a data inventory table is a simple way to keep track of the data associated with a project. KU Leuven recommends documenting information such as the following:
| Field |
Description |
| Name or identifier |
Assign a unique name or identifier to the dataset |
| Description |
Briefly describe the contents of the dataset |
| New or reused |
Indicate whether the data were generated for the project are being reused from another source |
| Ownership |
Indicate who holds the copyright or other relevant rights to the dataset |
| Type |
Specify the type of data contained in the dataset |
| Format / extension |
Specify the file format or extension |
| Volume |
Indicate the physical quantity or digital size of the data. |
| Protection |
Indicate whether the dataset contains personal or other sensitive data and, where applicable, what safeguards or processing measures are in place |
| Retention |
Indicate what will happen to the data after the project ends |
| Sharing |
Indicate whether the data will be shared and, where relevant, under what conditions |
To learn more
Training at UNIGE
12 June 2026