Gantt Chart for Clinical Data Management

Plan clinical data management timelines with a Gantt chart — covering CRF design, database build, UAT, database lock, SDTM/ADaM programming, and regulatory submission.

Gantt Chart for Clinical Data Management

Clinical data management (CDM) sits at the intersection of clinical operations and regulatory submission — responsible for converting raw patient data from clinical trial sites into clean, locked, analyzable datasets that support a regulatory filing. A Gantt chart for CDM is not just a project management tool; it is a quality control document that ensures ICH E6(R2) GCP compliance is built into the timeline, not bolted on at the end.

This guide covers the complete CDM timeline within a clinical trial, from protocol finalization through regulatory submission, including the milestone dependencies between CDM, biostatistics, and clinical operations that determine whether a study delivers on time.

CDM in Context: Where It Sits in the Trial Timeline

Clinical data management does not operate in isolation. The CDM timeline is deeply interdependent with:

A Gantt chart for CDM must show these dependencies explicitly, or the CDM team will be chronically surprised by schedule changes from other functions.

Phase 1: Protocol Finalization and CRF Design

Protocol finalization is the first upstream dependency for CDM. The CRF (Case Report Form) cannot be designed until the protocol specifies:

CRF design: The CDM team develops the CRF in parallel with the final protocol review, but CRF development cannot be completed until the protocol is locked. CRF design typically takes 4–8 weeks for a Phase II or III trial with 100–200 data points per visit.

CRF review by clinical team: Protocol-specified assessors (clinicians, biostatisticians, regulatory affairs) review the CRF draft for completeness and scientific accuracy. Budget 2 rounds of review at 1–2 weeks each.

CRF finalization and sign-off: The CDM manager, sponsor medical monitor, and statistician must formally sign off on the CRF before database build begins. This sign-off is a hard milestone — changes to the CRF after database build begins require database amendments that are costly and time-consuming.

Phase 2: Database Design and Build

Database design: Based on the finalized CRF, the database programmer designs the data structure in the EDC (Electronic Data Capture) system. Common EDC platforms: Medidata Rave/Veeva Vault CDMS, Oracle Clinical One, OpenClinica, Castor.

Database design includes:

Database build duration: 4–10 weeks for a moderately complex Phase II/III database.

Edit check programming: Edit checks are automated data validation rules that identify inconsistencies or out-of-range values when data is entered. Edit check programming typically takes 2–4 weeks and runs partially in parallel with database structure build.

Integration testing (IT): Before user acceptance testing, the database team performs internal integration testing to verify that all edit checks fire correctly and that all form logic works as designed.

Phase 3: User Acceptance Testing (UAT)

UAT is the formal verification that the database meets the CRF design specification and that all edit checks function correctly. UAT must be performed by testers who did not build the database — typically the CDM lead, a clinical scientist, and a biostatistician.

UAT test scripts: Written test cases that define what data to enter and what outcome to expect from each edit check. Test script development: 1–2 weeks.

UAT execution: Testers execute the scripts and document any discrepancies. Duration: 1–3 weeks.

UAT defect resolution: Programming team resolves defects and performs regression testing. A complex database may require 2–3 UAT rounds before all defects are resolved.

UAT sign-off: Formal sponsor approval that the database is fit for use in the study. This milestone is the hard predecessor to database go-live.

ICH E6(R2) requires that all system validation documentation (including UAT records) be maintained in the Trial Master File (eTMF). The UAT sign-off document is a GCP record.

Phase 4: Database Go-Live and Site Activation

Database go-live (production release): The validated database is released to study sites. This is sometimes called "database open" and marks the point at which live patient data can be entered.

Site activation and database go-live must be coordinated. Sites should not be activated (i.e., begin enrolling patients) before the database is available. Conversely, a database that sits live for months before site activation accumulates without data entry practice or training for site staff.

Site training: CDM or the sponsor's clinical operations team trains site coordinators on the EDC system. Training coordination takes 2–4 weeks for multi-site studies. Show training completion as a prerequisite to each site's first patient enrollment.

Phase 5: Data Collection and Ongoing Data Management

Data entry and ongoing cleaning: During the data collection phase (First Patient First Visit through Last Patient Last Visit), CDM performs ongoing data cleaning:

Ongoing cleaning is not a single task — it is a rolling process that should be tracked by query aging (the percentage of open queries that have been outstanding more than 30 days, 60 days, 90 days).

Interim analysis (if applicable): Many Phase II and Phase III trials have pre-specified interim analyses that require a partial database lock. Show interim analysis data cut dates as milestones, with the data cleaning sprint that precedes each cut.

Phase 6: End of Data Collection and Final Cleaning Sprint

Last Patient Last Visit (LPLV): The final study visit for the last patient enrolled. This is a major milestone that triggers the final data cleaning sprint.

Final cleaning sprint: After LPLV, CDM performs intensive final data cleaning:

The duration of the final cleaning sprint depends on how well ongoing cleaning was maintained. Studies with excellent ongoing cleaning (query aging kept below 30 days throughout) can close in 4–6 weeks. Studies with poor ongoing data quality may require 12–16 weeks for final cleaning.

Database lock review: Before database lock, a final quality review is conducted — often including a CRF completeness check, query closure audit, and review of all protocol deviations.

Phase 7: Database Lock

Database lock is the moment when no further changes are permitted to the clinical data. This is one of the most significant milestones in a clinical trial — it is the data quality commitment on which the regulatory submission will be based.

Database lock procedures typically require:

Database lock enables two parallel tracks: SAS programming/statistical analysis, and study report writing.

Phase 8: CDISC Dataset Production and Statistical Analysis

SDTM programming: CDISC Study Data Tabulation Model (SDTM) datasets are the standardized format required by FDA and EMA for regulatory submissions. SDTM programming begins immediately after database lock. Duration: 4–8 weeks for Phase II datasets; 8–16 weeks for Phase III.

ADaM programming: Analysis Data Model (ADaM) datasets derive from SDTM and are the datasets used by statisticians to produce tables, listings, and figures (TLFs). ADaM programming: 4–10 weeks.

Statistical analysis: With ADaM datasets available, biostatisticians execute the statistical analysis plan and produce TLFs. Duration: 4–8 weeks.

QC (quality control) of SAS programs: Each SDTM domain and ADaM dataset must be independently quality-controlled by a second programmer. Build QC time (typically 30–40% of programming time) into the Gantt chart.

Phase 9: Clinical Study Report and Regulatory Submission

Clinical Study Report (CSR): The CSR is the integrated document that presents the trial design, methods, results, and conclusions in a format specified by ICH E3. CSR writing typically takes 8–16 weeks and requires inputs from medical writing, biostatistics, CDM, and clinical operations.

Key regulatory submission milestones:

The data package for regulatory submission includes the SDTM datasets, ADaM datasets, analysis programs, dataset specifications, annotated CRF, data reviewer's guide, and define.xml file. CDM is responsible for all of these documents.

Building the CDM Gantt Chart

Structure the CDM Gantt chart with these parallel tracks:

  1. Setup track: Protocol → CRF design → database build → UAT → go-live
  2. Data collection track: Site activation → FPFV → ongoing cleaning → LPLV → final cleaning sprint
  3. Database lock and programming track: Lock → SDTM/ADaM programming → QC → statistical analysis
  4. Reporting track: CSR writing → regulatory submission

Key milestones with hard dependencies:

Use a free online Gantt chart maker to build this schedule at study startup and update it at each study milestone. Share the CDM Gantt chart with the full cross-functional team — when the clinical operations team understands that a 2-week delay in LPLV directly shifts the database lock and regulatory submission date, they manage the study differently. That shared schedule visibility is what makes CDM a strategic function rather than a bottleneck at the end of the trial.