Capture exports clean wide CSV or Excel files with a data dictionary, raw and decoded versions, and SDTM datasets as SAS XPT. Every statistics package reads them. Below is the code to read each file in, so you can test it before your first participant.
Free sandbox · No credit card · 21 CFR Part 11 aligned
| SUBJ | ECGPERF_CODE | ECGPERF_LABEL |
|---|---|---|
| 01-001 | Y | Yes |
| 01-002 | ND | Not done |
What you get, and what you do with it
What is in the export
The wide export has one row per form submission and one column per field. Dates are written as YYYY-MM-DD, and partial dates are preserved rather than guessed, so a date known only to the month stays a month. Not-done and not-applicable answers carry markers (ND, NA, VND and a separate not-applicable marker), and an empty cell means the answer is genuinely missing. That distinction matters: a skipped assessment and a forgotten one are different things in an analysis plan.
Raw and decoded versions exist because analysts and reviewers want different things. In the raw file a single-choice answer is its code, which is what you model. In the decoded file it is the label, which is what a medical monitor reads in a listing. The data dictionary ties them together, and every export can be limited by date range for interim snapshots or by site for site-level analyses.
Every export is recorded in an append-only export log with user, role, time, filters, row count and whether it was blinded. Blinded roles receive exports with the treatment arm withheld, as described on the blinded and unblinded roles page. If you need the full picture of review, freeze and lock before an export, see clinical data management.
Exports this month
6
Blinded exports
5
Rows in last export
1,284
Reading the CSV
File names are placeholders; use whatever your export is called. Reading everything as text first lets you see the ND, NA and VND markers before you decide how to recode them.
| Tool | Code |
|---|---|
| R (readr) | library(readr); vs <- read_csv("vital_signs.csv", na = "", col_types = cols(.default = col_character())) |
| SAS | proc import datafile="vital_signs.csv" out=work.vs dbms=csv replace; guessingrows=max; run; |
| SPSS | File > Import Data > CSV Data, or the Text Import Wizard, then paste the generated GET DATA syntax into your syntax file so the import is repeatable. |
| Stata | import delimited "vital_signs.csv", varnames(1) encoding("UTF-8") stringcols(_all) clear |
| Python (pandas) | import pandas as pd; vs = pd.read_csv("vital_signs.csv", dtype=str, keep_default_na=False) |
Why na = "" in readr and keep_default_na=False in pandas: both libraries treat the text NA as missing by default, which would merge "not available" with "missing". Keeping it as text lets you recode on purpose.
Reading SDTM XPT
Each SDTM domain is one XPT file. These are the standard calls from each package's documentation.
| Tool | Code |
|---|---|
| R (haven) | library(haven); dm <- read_xpt("dm.xpt") |
| SAS | libname xpt xport "dm.xpt"; proc copy in=xpt out=work; run; |
| SPSS | GET SAS DATA='dm.xpt'. |
| Stata (16 or later) | import sasxport5 "dm.xpt", clear |
| Python (pandas) | dm = pd.read_sas("dm.xpt", format="xport", encoding="utf-8") |
Readers differ in which transport version they accept. The SAS xport engine, Stata sasxport5 and pandas read version 5 files; haven reads versions 5 and 8. Open one file from a sandbox export with your own tool before you build a pipeline around it.
Practical details
The first job after reading a wide export is to decide what ND, NA and VND mean for each analysis. In R, a pattern that works is to read as text, then convert deliberately: replace the markers with NA in the numeric columns you are analysing, keep a separate indicator column where "not done" is itself informative, and convert with as.numeric. In Stata, import delimited will make a column string if a marker appears in it, so use the stringcols option shown above and destring after you have recoded.
Partial dates survive the export, so do not rely on a plain date conversion. In R, the lubridate package handles them: ymd(x, truncated = 2) parses a full date, a year and month, or a year alone, returning the first day of the period, so decide in your statistical analysis plan how partial dates are imputed before you use them. In pandas, pd.to_datetime(x, errors="coerce") returns NaT for partial values, which is a prompt to handle them explicitly instead of silently losing them.
Labels come from the data dictionary. Haven keeps variable labels on columns it reads from XPT as attributes, and you can read them with attr(dm$AGE, "label"). For CSV files, join the dictionary to your column names to build labels and factor levels once, then reuse the script for every export so each interim snapshot is analysed the same way.
If your team works in SPSS or Stata and wants native files, read the CSV or XPT in R and write them out. haven::write_sav(vs, "vital_signs.sav") writes an SPSS file and haven::write_dta(vs, "vital_signs.dta") writes a Stata file, carrying across any variable labels you have set. This runs on your side after export, and it is repeatable, so the conversion can be part of the analysis script.
Build a form in the free sandbox, enter a few sample subjects and export. Run the snippets above before the first real participant is screened.
Workflow
Check every endpoint component is a coded field, units are fixed and "not done" is separate from missing. The cheapest fixes happen before the first participant.
Enter sample data, export raw and decoded files with the data dictionary, and note the exact file layout.
Use the snippets above in R, SAS, SPSS, Stata or Python, and store it with the analysis code.
For interim snapshots, filter by date range, freeze or lock the data, export, and re-run the same script. Record the export log entry in your data management file.
Blinded roles receive exports without treatment arm. Plan how any unblinded outputs, for example for a data monitoring committee, will be produced by people with the right role.
Choosing a file
| Job | Use | Why |
|---|---|---|
| Modelling and tables | Raw wide CSV with data dictionary | Coded values, stable column names |
| Data listings and review | Decoded wide CSV or Excel | Labels instead of codes, readable by a clinician |
| Site-level or interim analysis | Either, with site or date-range filter | Snapshot matches the cut-off |
| Regulatory submission datasets | SDTM XPT with Define-XML | CDISC domains, one record per observation |
| Native SPSS or Stata files | Read CSV or XPT in R, then write_sav or write_dta | No native writer in the platform, one extra step |
Export the wide CSV or Excel file with its data dictionary, then read it with readr::read_csv or readxl. For SDTM datasets, read the SAS XPT files with haven::read_xpt. Example code for each is on this page.
Capture exports CSV, Excel and SDTM SAS XPT files. It does not write native .sas7bdat or .sav files. SAS reads CSV with PROC IMPORT and XPT with the XPORT engine, SPSS reads CSV through the import dialog and XPT with GET SAS, and Stata uses import delimited or import sasxport5.
Raw files carry coded values, which is what you model. Decoded files carry labels, which is what reviewers read in listings. The data dictionary links the two.
With markers: ND for not done, NA for not available, VND for visit not done and a separate not-applicable marker. An empty cell means the answer is missing.
Yes. Exports can be filtered by date range and by site, which supports interim snapshots and site-level analyses.
Yes. Each export goes into a separate append-only export log with user, role, time, filters, row count and whether it was blinded.
No. The sandbox is free with every feature, no credit card and no time limit. You pay only once you go live with real participants, so the export can be tested end to end beforehand.
Keep exploring
CDISC SDTM export
SAS XPT datasets with Define-XML.
EDC for biostatisticians and SAS programmers
Role page for the analysis team.
Clinical data management software
Review, freeze, lock and export.
Capture for academia
Investigator-led studies.
Blinded and unblinded role management
Blinded exports stay blinded.
Build a form, enter sample data and test your R, SAS, SPSS or Stata read-in on a real export. No credit card, no time limit.