Zubair Khalid

Virologist/Molecular Biologist | Veterinarian | Bioinformatician

Conventional & Molecular Virology • Vaccine Development • Computational Biology

Dr. Zubair Khalid is a veterinarian and virologist specializing in conventional and molecular virology, vaccine development, and computational biology. Dedicated to advancing animal health through innovative research and multi-omics approaches.

Dr. Zubair Khalid - Veterinarian, Virologist, and Vaccine Development Researcher specializing in Computational Biology, Multi-omics, Animal Health, and Infectious Disease Research

Section: Infrastructure, Cloud & Policy

Data Stewardship in Applied Research: Roles, Responsibilities, and Best Practices

Data stewardship in applied research is the active management of research data throughout its lifecycle to ensure it remains findable, accessible, interoperable, and reusable. For students, researchers, analysts, and life-science professionals, a data steward acts as the designated person who translates policy requirements into daily research practice, maintains data quality, and prepares datasets for sharing and long-term preservation. This article explains the data steward role using established programs and frameworks, then provides a practical implementation framework, a role description template, and a checklist for embedding stewardship into research projects.

What Data Stewards Actually Do in Research Settings

A data steward is not an IT administrator or a lab manager, though the role overlaps with both. The steward is the person who owns the practical responsibility for how data are collected, documented, stored, protected, and shared. In applied research, this means making concrete decisions about file naming, metadata standards, version control, access permissions, and repository selection before data collection begins.

The FAIR Guiding Principles formalize what good stewardship looks like. Published in Scientific Data, these principles state that data should be findable, accessible, interoperable, and reusable. The PubMed record of the FAIR principles explains that the intent is to enhance the ability of machines to automatically find and use data, in addition to supporting reuse by individuals. This machine-readability requirement is a core distinction of modern data stewardship. A steward must ensure that metadata schemas, file formats, and data dictionaries are structured so that both humans and software can interpret them.

In practice, the steward role includes several recurring tasks. The steward writes and maintains the data management plan, enforces naming conventions, tracks dataset versions, documents processing steps, manages access controls, and prepares data packages for repository deposit. The steward also trains team members on these practices and audits compliance. A national survey of data stewardship in Italy found that most institutions employ staff performing data management support tasks that map to the definition of data stewardship, though these professionals are rarely formally recognized with that title. The most common tasks include FAIR data management support, data management plan writing, researcher training, and policy consultancy. This finding, reported in Mapping data stewardship in Italy, shows that the work happens widely even where the job title does not exist.

Why Applied Research Projects Need a Designated Steward

Applied research projects generate heterogeneous data from multiple sources. A single project may combine laboratory measurements, field observations, instrument outputs, survey responses, and administrative records. Each data type has different quality characteristics, documentation needs, and reuse potential. Without a designated steward, these responsibilities scatter across team members and typically fall to whoever happens to be available when a problem arises.

The consequences of poor stewardship are measurable. A methods paper on using administrative and surveillance databases in healthcare epidemiology and antimicrobial stewardship research notes that data quality issues exist and require careful consideration and potential validation of data. The authors provide a checklist to help aid study development when using administrative data. Their guidance, published in Research Methods in Healthcare Epidemiology and Antimicrobial Stewardship, applies broadly to any research that relies on secondary data sources. The steward is the person who applies that checklist and documents the validation steps.

Artificial intelligence and machine learning workflows have raised the stakes for data stewardship. A Nature article on scientific discovery in the age of artificial intelligence states that challenges posed by poor data quality and stewardship remain critical areas of focus for AI innovation. Models trained on poorly documented or inconsistently labeled data produce unreliable outputs, and those outputs are difficult to trace back to their data origins. The steward role becomes essential for maintaining the provenance records that make AI-assisted research auditable.

At a Glance: Data Stewardship Roles and Responsibilities

Responsibility Area Typical Tasks Common Failure Without Stewardship
Data planning Write data management plans, define metadata standards, select repositories No documentation of data origins or processing steps
Data quality Validate entries, track errors, document corrections, monitor completeness Undetected errors propagate into analysis and published results
Data access and security Manage permissions, enforce consent terms, control sharing Unauthorized access or sharing that violates participant agreements
Data preservation Format files for longevity, deposit in repositories, assign identifiers Data lost when staff leave or storage media fail
Training and compliance Train team members, audit practices, respond to policy changes Inconsistent practices across team members and projects

Core Principles of Research Data Stewardship

FAIR Principles as the Operating Standard

The FAIR principles provide the operational standard for data stewardship. Findable means data and metadata have unique identifiers and are described in searchable registries. Accessible means data can be retrieved using standard protocols, with clear authentication and authorization procedures. Interoperable means data use formal, shared vocabularies and reference other data. Reusable means data are described with rich metadata and meet domain-relevant community standards.

The FAIR Guiding Principles were designed to act as a guideline for those wishing to enhance the reusability of their data holdings. The PubMed record emphasizes that the principles put specific emphasis on enhancing the ability of machines to automatically find and use data. For applied researchers, this machine emphasis means that human-readable documentation alone is insufficient. Metadata must be structured in machine-readable formats such as JSON, XML, or domain-specific schemas.

Data Sharing Policies and Compliance

Research funders and institutions increasingly mandate data sharing. The NIH Genomic Data Sharing Policy establishes expectations for sharing genomic data generated through NIH-funded research. The steward must understand which policies apply to a given project and build compliance into the data management plan from the start. This includes knowing what data must be shared, through which repositories, and under what access conditions.

Compliance is not a one-time event. Policies change, and the steward must track those changes and adjust project practices accordingly. The steward also serves as the point of contact when questions arise about whether a particular data use is permitted under a consent agreement or data use limitation.

Ethical Stewardship Beyond Technical Management

Data stewardship carries ethical dimensions that go beyond file management. A 2024 study on data stewardship in FTLD research examined investigator and participant views on data stewardship practices in frontotemporal lobar degeneration research. The study identified three meta themes: perspectives on data sharing, experiences with enrollment and participation, and data management and security as mechanisms for participant protections. The results offer initial insights on ethical challenges to data stewardship aimed at informing future guidelines and policies.

The ethical dimension means the steward must understand what participants were told during consent, what data use limitations apply, and how to honor those commitments during sharing and reuse. The steward also handles the return of individual results to participants, a practice that federal policies and guidelines have expanded. This requires coordination with ethics boards, clinical teams, and data repositories.

Building a Data Stewardship Program for Your Research Group

Step 1: Define the Steward Role and Scope

Start by writing a role description that fits your research context. The role description template below provides a starting point. Adapt it to your institution, funding sources, and data types.

Role Description Template: Research Data Steward

Position purpose: The research data steward manages research data throughout the project lifecycle, ensuring data quality, documentation, security, and compliance with institutional and funder policies.

Core responsibilities:

  • Develop and maintain the project data management plan
  • Define and enforce file naming, versioning, and metadata standards
  • Conduct data quality checks and document corrections
  • Manage data access permissions and security controls
  • Prepare data packages for repository deposit and sharing
  • Train research team members on data management practices
  • Monitor compliance with funder, institutional, and ethical requirements
  • Serve as the point of contact for data-related inquiries

Required competencies:

  • Knowledge of research data management practices and FAIR principles
  • Familiarity with domain-specific data formats and metadata standards
  • Understanding of data sharing policies and consent requirements
  • Ability to document workflows clearly and train others
  • Attention to detail and systematic approach to quality control

Reporting structure: The steward reports to the principal investigator or project lead and coordinates with institutional research support services.

Step 2: Assess Current Data Practices

Before implementing new practices, assess what currently happens with research data in your group. Review recent projects and answer these questions:

  • Where are data stored during active analysis?
  • What file naming conventions are in use, if any?
  • How are data versions tracked?
  • What metadata are recorded, and where?
  • Who has access to which data, and how is access controlled?
  • What happens to data when a team member leaves?
  • Which repositories or archives hold completed datasets?

This assessment identifies gaps and priorities. A group that loses data when students graduate has a preservation problem. A group that cannot reproduce its own analyses has a documentation problem. A group that shares data without checking consent terms has a compliance problem.

Step 3: Write the Data Management Plan

The data management plan is the central document of the stewardship program. It should specify:

  • What data will be collected or generated
  • What metadata standards will be used
  • How data will be stored and backed up during the project
  • Who has access and under what conditions
  • How data quality will be maintained
  • What data will be shared and through which repositories
  • How long data will be preserved
  • Who is responsible for each stewardship task

The plan should be written at project start and updated when significant changes occur. The steward owns the plan but should consult with the principal investigator, ethics board, and institutional research office.

Step 4: Implement Standards and Tools

Select standards and tools that match your research domain. For bioinformatics and life-science research, relevant resources include the EMBL-EBI Training materials and the NCBI Data Resources. These platforms provide domain-specific guidance on data formats, submission requirements, and repository options.

Choose tools that the team will actually use. A sophisticated electronic lab notebook that nobody opens provides less value than a simple shared folder with enforced naming conventions. The steward should select tools based on team capacity and willingness to adopt them.

Step 5: Train the Team

Training is a core stewardship task. The steward trains new team members on data management expectations and provides refresher sessions when practices change. Training should cover the data management plan, file naming and versioning rules, metadata documentation requirements, and data sharing procedures.

The research data management and data stewardship competences in university curriculum literature shows that these skills are increasingly taught in formal education settings. However, most researchers still learn data management on the job. The steward fills this gap by providing project-specific training that connects general principles to concrete practices.

Step 6: Monitor and Audit

Stewardship requires ongoing monitoring, beyond planning. The steward should conduct periodic audits of project data to check compliance with the data management plan. Audits examine file organization, metadata completeness, version control, and access permissions. Findings from audits feed back into training and plan updates.

Practical Implementation Checklist for Research Projects

Use this checklist when starting a new project or when taking over stewardship of an existing one.

Planning phase

  • Designate a data steward for the project
  • Write the data management plan
  • Identify applicable funder and institutional data policies
  • Select metadata standards and file formats
  • Define file naming and versioning conventions
  • Establish storage and backup locations
  • Document consent terms and data use limitations

Active research phase

  • Enforce naming and versioning conventions
  • Record metadata at time of data collection
  • Conduct regular data quality checks
  • Document all data processing and analysis steps
  • Manage access permissions and security
  • Track data use and sharing requests
  • Update the data management plan when methods change

Completion and sharing phase

  • Validate data completeness and quality
  • Prepare metadata and data dictionaries
  • Select appropriate repositories
  • Deposit data packages with persistent identifiers
  • Confirm compliance with consent and policy requirements
  • Document repository locations for publications
  • Plan long-term preservation and access

Records and Measurements for Stewardship Accountability

Data stewardship should be documented and measurable. The steward maintains records that demonstrate what was done, when, and by whom. These records serve multiple purposes: they support reproducibility, provide evidence for compliance audits, and enable continuity when staff change.

Essential records include:

  • The data management plan with version history
  • File naming and metadata standards documentation
  • Data quality check logs showing what was checked and what was found
  • Correction logs documenting changes to data and rationale
  • Access control records showing who had access and when
  • Data sharing logs tracking requests and approvals
  • Repository deposit records with persistent identifiers
  • Training attendance and materials

Measurement helps the steward identify problems before they become crises. Track metrics such as the percentage of datasets with complete metadata, the time between data collection and repository deposit, the number of data quality issues found during audits, and the proportion of team members who follow naming conventions. These metrics provide early warning of stewardship failures.

Common Failure Patterns in Research Data Stewardship

Failure Pattern 1: Stewardship as an Afterthought

Many projects treat data management as something to address at the end, when writing up results or preparing for repository deposit. This approach fails because data quality and documentation problems compound over time. Files with unclear names, undocumented processing steps, and missing metadata cannot be reliably reconstructed after the fact. The steward must be involved from project design onward.

Failure Pattern 2: The Steward as a Sole Point of Failure

When one person holds all data management knowledge, the project becomes vulnerable to that person's departure. The steward should document procedures so that others can take over, and should train multiple team members on core practices. Cross-training is a stewardship responsibility, not an optional extra.

Failure Pattern 3: Documentation That Nobody Reads

A data management plan that sits in a drawer provides no value. The steward must translate the plan into daily practices and make documentation accessible where work happens. This means integrating checklists into lab meetings, embedding metadata templates into data collection tools, and reviewing the plan with new team members.

Failure Pattern 4: Treating All Data the Same

Different data types require different stewardship approaches. Raw instrument outputs need different documentation than derived analysis files. Identifiable human data need different security controls than public reference data. The steward must tailor practices to data sensitivity, complexity, and reuse potential instead of applying one standard to everything.

Failure Pattern 5: Ignoring the Human Dimensions

A 2021 article on stewardship versus interdisciplinarity highlights tensions that arise when stewardship requirements meet disciplinary norms. Researchers trained in different fields have different expectations about data sharing, documentation, and authorship. The steward must navigate these differences through communication and negotiation, beyond policy enforcement.

Limitations and Boundaries of the Steward Role

The data steward role has limits that should be acknowledged from the start. The steward does not replace the principal investigator's scientific judgment about what data mean or how they should be interpreted. The steward does not make decisions about research direction or hypothesis testing. The steward does not override consent terms or institutional policies based on personal judgment.

The steward also operates within resource constraints. A survey of the infectious diseases and antimicrobial stewardship pharmacist workforce in the United States found that respondents frequently indicated they lacked adequate job resources. While that survey addressed a different stewardship domain, the finding about resource limitations applies broadly. Data stewardship requires dedicated time, appropriate tools, and institutional support. A steward assigned stewardship duties on top of a full research workload will struggle to maintain quality.

The National Academies report on ensuring the integrity, accessibility, and stewardship of research data in the digital age addresses the structural challenges of data stewardship at the institutional level. Individual stewards cannot solve systemic problems such as inadequate infrastructure or unclear institutional policies. The steward should escalate systemic issues to institutional leadership instead of attempting to work around them.

Professional Escalation Criteria for Data Stewards

The steward should escalate concerns when they encounter situations beyond their authority or capacity. Clear escalation criteria protect both the steward and the research project.

Escalate to the principal investigator when:

  • Data quality issues could affect the validity of research findings
  • A data sharing request raises consent or policy questions
  • A team member repeatedly fails to follow data management requirements
  • Storage or infrastructure capacity threatens data preservation
  • A data breach or unauthorized access is suspected

Escalate to the institutional research office or ethics board when:

  • A data sharing request conflicts with consent terms
  • A funder or publisher requirement cannot be met with current resources
  • A policy interpretation is unclear or contested
  • A compliance audit identifies systemic problems

Escalate to institutional leadership when:

  • Data infrastructure is inadequate for the research program
  • Stewardship responsibilities are not resourced adequately
  • Institutional policies conflict with funder requirements
  • Systemic data management problems affect multiple projects

The steward should document escalation decisions and outcomes. This documentation provides a record of due diligence and supports future planning.

Data Stewardship in the Age of Machine Learning and AI

The integration of artificial intelligence into scientific discovery has changed data stewardship requirements. The Nature article on scientific discovery in the age of artificial intelligence notes that AI methods help scientists generate hypotheses, design experiments, collect and interpret large datasets, and gain insights that might not have been possible using traditional methods. However, the article also states that challenges posed by poor data quality and stewardship remain critical areas of focus.

For the data steward, AI workflows introduce specific requirements. Training datasets must be documented with sufficient detail to support model interpretation and debugging. Data provenance must be tracked so that model outputs can be traced back to their inputs. Version control becomes more complex when both data and models change over time. The steward must ensure that datasets used for training and validation are properly separated and documented.

The practical guide to bioimaging research data management in core facilities illustrates how stewardship requirements scale with data complexity. Bioimage data are generated in diverse research fields throughout the life and biomedical sciences, and their potential for advancing scientific progress via modern data-driven discovery approaches reaches beyond disciplinary borders. The guide notes that implementing the FAIR principles into daily routines is an essential but challenging task for researchers and research infrastructures. Imaging core facilities are positioned at the intersection of research groups, IT infrastructure providers, institutional administration, and microscope vendors, making them natural leaders in this transformation.

Building Institutional Support for Data Stewardship

Individual stewards are most effective when institutions support their work. Institutional support takes several forms: formal recognition of the steward role, dedicated time for stewardship activities, access to training and professional development, and investment in data infrastructure.

The Italian national survey found that data stewards are rarely formally recognized with that title, and the variety of job titles observed reflects the absence of a standardized professional profile. This lack of recognition creates problems. Without a formal title, stewards may lack the authority to enforce data management requirements. Without a defined career path, institutions struggle to recruit and retain skilled stewards.

Institutions can address these problems by creating formal steward positions, including stewardship responsibilities in job descriptions, and providing advancement opportunities. Institutions should also invest in training programs. The E-infrastructures Austria training seminar for research data stewardship represents one model of structured professional development for data stewards. Similar programs help stewards build the technical and interpersonal skills the role requires.

Data Stewardship Across Research Domains

While the core principles of stewardship apply broadly, implementation varies by research domain. The steward must understand the data conventions, repositories, and quality standards of their specific field.

In genomics and bioinformatics, stewardship involves managing sequence data, variant calls, and phenotype data. The NCBI Data Resources provide repositories and tools for these data types. The NIH Genomic Data Sharing Policy establishes expectations for data sharing and access controls. The steward must understand submission formats, metadata requirements, and access tiers.

In clinical and health research, stewardship involves protecting participant privacy while enabling data sharing. The data stewardship in FTLD research study found that data management and security serve as mechanisms for participant protections. The steward must balance the scientific value of data sharing against the ethical obligations to research participants.

In imaging and microscopy, stewardship involves managing large, complex datasets with specialized metadata requirements. The bioimaging data management guide emphasizes that making bioimaging data FAIR is essential for remaining competitive and at the forefront of research. The steward must work with core facilities and vendors to ensure that instrument outputs include the metadata needed for future interpretation.

In research using administrative and surveillance data, stewardship involves validating secondary data sources and documenting their limitations. The healthcare epidemiology methods paper provides a checklist for this work. The steward must understand the data generation process, including how the data were collected, coded, and cleaned, to assess their fitness for research use.

Frequently Asked Questions

What is the difference between a data steward and a data manager?

A data manager typically handles the technical infrastructure of data storage, backup, and retrieval. A data steward has broader responsibility for data quality, documentation, policy compliance, and ethical use. The steward decides what standards to apply and ensures the team follows them, while the manager operates the systems that store and serve the data. In small research groups, one person often fills both roles.

Do I need a formal data steward title to perform stewardship duties?

No. The Italian national survey found that most institutions employ staff performing data management support tasks that map to the definition of data stewardship, though these professionals are rarely formally recognized with that title. What matters is that someone has explicit responsibility for data stewardship tasks and the authority to enforce data management practices. A formal title helps with authority and career recognition but is not a prerequisite for doing the work.

How much time does data stewardship require?

Time requirements vary with project complexity, data volume, and team size. A small project with a few datasets and a stable team may require only a few hours per month. A large multi-site project with heterogeneous data types and complex sharing requirements may require a dedicated full-time steward. The steward should track time spent on stewardship tasks to inform project planning and resource requests.

What training do I need to become a data steward?

Training needs depend on your background and research domain. Core competencies include research data management practices, FAIR principles, metadata standards, data sharing policies, and data quality methods. Domain-specific training covers the repositories, formats, and standards of your research field. Resources such as EMBL-EBI Training provide relevant materials for life-science researchers. Many institutions and professional organizations offer data stewardship workshops and courses.

How do I handle data sharing requests that raise consent concerns?

When a data sharing request raises consent or policy questions, escalate to the principal investigator and the institutional research office or ethics board. The steward should not make independent judgments about whether a proposed use falls within the original consent terms. Document the request, the concerns, and the decision. The NIH Genomic Data Sharing Policy provides guidance on data use limitations for genomic data, and similar frameworks apply in other domains.

What should I do if I discover data quality problems in a completed analysis?

If data quality issues could affect the validity of research findings, escalate immediately to the principal investigator. Document what was found, when it was found, and what data are affected. The steward should work with the research team to assess the impact on published or in-progress results and to determine whether corrections or retractions are needed. The healthcare epidemiology methods paper emphasizes the importance of validation when using secondary data, and the same principle applies to primary data quality problems.

How do I choose a repository for data deposit?

Repository selection depends on your research domain, data type, and funder requirements. Domain-specific repositories often provide the best fit because they understand the data formats and metadata standards of the field. The NCBI Data Resources serve genomics and related life-science data. General repositories may be appropriate for data types without a domain-specific home. The steward should verify that the chosen repository meets funder and publisher requirements and provides persistent identifiers for deposited datasets.

What happens to research data when a project ends?

The data management plan should specify what happens to data at project end. Typically, data are deposited in a repository for long-term preservation and sharing, with access controls applied as needed. The steward ensures that the deposit package includes complete metadata, documentation, and any code or workflows needed to interpret the data. The steward also confirms that the deposit complies with consent terms and funder requirements before transfer.

Related Bioinformatics Guides

References and Further Reading

This article is educational and does not replace validated analysis plans, institutional policy, clinical interpretation, or specialist review.