Skip to content

CDISC: SDTM, ADaM, Define-XML

Sources, the mapping specification, the build, checks, licences and limits.

The CDISC pipeline turns an export from any EDC into the files a sponsor, a CRO or a regulator expects: the source as CDISC ODM, Excel and SAS, then SDTM and ADaM datasets with Define-XML 2.1. Every step is driven by a mapping specification that you can read, edit and keep. No AI is used: the same specification and data always give the same result.

1. Upload the export

Accepted: one ODM 1.3 XML file, or CSV/XLSX files with one file or sheet per form (a "modular" export), or one flat file, or a ZIP of them. Key columns are recognised by name:

  • subject: SUBJID, Subject, Patient, Screening number…;
  • site: SITEID, Site, Centre;
  • visit: Visit, Folder, Event, InstanceName, redcap_event_name;
  • line of a log form: Line, RecordPosition, AESEQ, redcap_repeat_instance.

Every other column becomes an item; its data type (integer, float, date, text) is taken from the values. Dates like 05.03.2024, 2024-03-05 or 05MAR2024 are recognised. System columns such as "last modified" are skipped.

Right after the upload, the Source tab gives the data back as CDISC ODM 1.3.2 (metadata and clinical data in one XML), Excel (a sheet per table) and SAS (a transport file per table plus import.sas, which applies the code lists as SAS formats). This alone converts a modular or flat export to ODM.

Upload pseudonymised data only. No names, phone numbers or document numbers. The source and the last package are stored in your account until you delete the study; nobody else can open them.

2. Check the specification

The specification is proposed from the data. Columns named the CDASH way (AETERM, AESTDAT, VSORRES, SYSBP…) map straight to SDTM; tables are assigned to domains by name (AE, "Adverse events", "НЯ"…). It has these parts:

  • Settings: STUDYID, the USUBJID pattern, COUNTRY, where the planned arm, the visit date, the informed consent and randomisation dates come from, the treatment-emergent window and age groups;
  • Domains: which source tables feed DM, AE, CM, MH, EX, DS, SV, VS and LB, plus the trial design datasets TA, TE, TV and TS;
  • Variables: a rule for each SDTM variable (copy a column, ISO 8601 date, constant, code list, earliest or latest date, ongoing, SUPP--, derived, empty);
  • Tests: wide findings columns (one column per test, as most eCRFs collect vital signs) become one record per test with the CDISC test code and unit;
  • Maps: collected value to CDISC submission value per code list, for example "Лёгкая" to MILD. Empty rows are the ones to fill in;
  • Visits, Arms, Trial: VISITNUM, planned study day and epoch, ARMCD, trial summary parameters.

Settings, domains, maps and trial parameters can be edited on the page. For the rest, download the specification as Excel, edit it and upload it back. "Propose again from the data" starts over and keeps the country and the trial parameters.

3. Build

One build produces the whole package. It stops with a list when the specification refers to a table or column that does not exist. Derived automatically: USUBJID, --SEQ, study days, VISITNUM and VISIT, EPOCH, the DM reference dates (RFSTDTC, RFXSTDTC…), ARM and ACTARM, screen failures (ARMNRS), baseline flags (--LOBXFL), normal range indicators, standard results and the trial design datasets.

ADaM: ADSL (one record per randomised or treated subject: TRT01P/A, TRTSDT/TRTEDT, RANDFL, ITTFL, SAFFL, AGEGR1, EOSSTT, baseline height, weight and BMI), ADAE (treatment-emergent flag and first occurrence flags), ADVS and ADLB (AVAL, BASE, CHG, PCHG, ADY). Each derivation is described in define.xml as a method.

4. What you download

  • sdtm/ and adam/: datasets as SAS V5 transport (.xpt), CDISC Dataset-JSON 1.1 and Excel, define.xml (Define-XML 2.1) with a readable define.html;
  • source/: the source as ODM, Excel and SAS;
  • spec/sdtm-spec.xlsx: the specification used for this build;
  • checks/checks.xlsx: the check report;
  • sas/import.sas: reads every .xpt into SAS WORK.

Checks

The checks (IDs YE-…) cover required variables, keys and --SEQ, ISO 8601 dates, start after end, controlled terminology, SAS V5 limits (names of 8 characters, labels of 40, values of 200 bytes), non-ASCII text, and the ADaM structure. They are our own checks, not CDISC CORE rules and not Pinnacle 21. Before a submission, run a conformance validator, for example the open-source CDISC CORE engine.

Limits of this version

  • Medical coding comes from your own dictionaries. AEDECOD, AEBODSYS and the MedDRA hierarchy (AE, MH) and CMDECOD with CMCLAS (CM, ATC) are filled from Medical coding: load your organisation's licensed MedDRA or a full ATC there and code the terms. The build uses your decisions and single exact matches only, names the dictionary and version in define.xml, and lists terms still to code. Without such a dictionary these variables stay empty unless your export has coded columns. WHODrug is not supported yet.
  • Values stay in the collected language. FDA expects ASCII text in transport files; Russian or French verbatim terms are reported as warnings.
  • No unit conversion and no imputation of partial dates.
  • Domains: DM, AE, CM, MH, EX, DS, SV, VS, LB, TA, TE, TV, TS and SUPP--. Other domains are not built yet.
  • Controlled terminology: a bundled subset covering these domains. Load the full release you use (the NCI EVS "SDTM Terminology.txt" file) on the module's main page.

The package is a draft for a statistical programmer to review, not a validated submission.

Standards and sources

© 2026 Т.Г. (contact@youedc.com), youEDC. All rights reserved. Terms of use. Data entered by users belongs to them.