Data exports · For statisticiansUpdated October 10, 2026

Export EDC data to R, SAS, SPSS and Stata

Capture exports clean wide CSV or Excel files with a data dictionary, raw and decoded versions, and SDTM datasets as SAS XPT. Every statistics package reads them. Below is the code to read each file in, so you can test it before your first participant.

  • CSV and Excel with data dictionary
  • Raw and decoded files
  • SDTM as SAS XPT

Free sandbox · No credit card · 21 CFR Part 11 aligned

Data and exports
Clinical data (codes + labels)CDISC SDTM + Define-XMLData dictionaryAnnotated CRFE2B safetyAudit trail
SUBJECGPERF_CODEECGPERF_LABEL
01-001YYes
01-002NDNot done

What you get, and what you do with it

  • CSV and Excel exports, wide format with one row per submission, filterable by date range and by site.
  • An automatic data dictionary with every variable, its type and its code list, so a program can apply labels and formats without guessing.
  • Raw files with coded values for analysis, and decoded files with labels for review and listings.
  • SDTM datasets as SAS XPT with Define-XML, for teams that work in CDISC domains. See the CDISC SDTM export page.
  • Capture does not write native .sas7bdat or .sav files and there is no API on this page. The route into SAS, SPSS and Stata is the CSV or XPT export, and the snippets below cover it, including R as a converter if you want .sav or .dta files.

What is in the export

Files an analyst can read without a phone call

The wide export has one row per form submission and one column per field. Dates are written as YYYY-MM-DD, and partial dates are preserved rather than guessed, so a date known only to the month stays a month. Not-done and not-applicable answers carry markers (ND, NA, VND and a separate not-applicable marker), and an empty cell means the answer is genuinely missing. That distinction matters: a skipped assessment and a forgotten one are different things in an analysis plan.

Raw and decoded versions exist because analysts and reviewers want different things. In the raw file a single-choice answer is its code, which is what you model. In the decoded file it is the label, which is what a medical monitor reads in a listing. The data dictionary ties them together, and every export can be limited by date range for interim snapshots or by site for site-level analyses.

Every export is recorded in an append-only export log with user, role, time, filters, row count and whether it was blinded. Blinded roles receive exports with the treatment arm withheld, as described on the blinded and unblinded roles page. If you need the full picture of review, freeze and lock before an export, see clinical data management.

Export log (demo study)

Exports this month

6

Blinded exports

5

Rows in last export

1,284

Raw wide CSV3/6
Decoded wide CSV2/6
SDTM XPT1/6
Illustrative numbers. Each real export is logged with user, role, time and filters.

Reading the CSV

Read a wide CSV export in five tools

File names are placeholders; use whatever your export is called. Reading everything as text first lets you see the ND, NA and VND markers before you decide how to recode them.

ToolCode
R (readr)library(readr); vs <- read_csv("vital_signs.csv", na = "", col_types = cols(.default = col_character()))
SASproc import datafile="vital_signs.csv" out=work.vs dbms=csv replace; guessingrows=max; run;
SPSSFile > Import Data > CSV Data, or the Text Import Wizard, then paste the generated GET DATA syntax into your syntax file so the import is repeatable.
Stataimport delimited "vital_signs.csv", varnames(1) encoding("UTF-8") stringcols(_all) clear
Python (pandas)import pandas as pd; vs = pd.read_csv("vital_signs.csv", dtype=str, keep_default_na=False)

Why na = "" in readr and keep_default_na=False in pandas: both libraries treat the text NA as missing by default, which would merge "not available" with "missing". Keeping it as text lets you recode on purpose.

Reading SDTM XPT

Read the SAS XPT datasets in each tool

Each SDTM domain is one XPT file. These are the standard calls from each package's documentation.

ToolCode
R (haven)library(haven); dm <- read_xpt("dm.xpt")
SASlibname xpt xport "dm.xpt"; proc copy in=xpt out=work; run;
SPSSGET SAS DATA='dm.xpt'.
Stata (16 or later)import sasxport5 "dm.xpt", clear
Python (pandas)dm = pd.read_sas("dm.xpt", format="xport", encoding="utf-8")

Readers differ in which transport version they accept. The SAS xport engine, Stata sasxport5 and pandas read version 5 files; haven reads versions 5 and 8. Open one file from a sandbox export with your own tool before you build a pipeline around it.

Practical details

Markers, dates and labels once the data is in

The first job after reading a wide export is to decide what ND, NA and VND mean for each analysis. In R, a pattern that works is to read as text, then convert deliberately: replace the markers with NA in the numeric columns you are analysing, keep a separate indicator column where "not done" is itself informative, and convert with as.numeric. In Stata, import delimited will make a column string if a marker appears in it, so use the stringcols option shown above and destring after you have recoded.

Partial dates survive the export, so do not rely on a plain date conversion. In R, the lubridate package handles them: ymd(x, truncated = 2) parses a full date, a year and month, or a year alone, returning the first day of the period, so decide in your statistical analysis plan how partial dates are imputed before you use them. In pandas, pd.to_datetime(x, errors="coerce") returns NaT for partial values, which is a prompt to handle them explicitly instead of silently losing them.

Labels come from the data dictionary. Haven keeps variable labels on columns it reads from XPT as attributes, and you can read them with attr(dm$AGE, "label"). For CSV files, join the dictionary to your column names to build labels and factor levels once, then reuse the script for every export so each interim snapshot is analysed the same way.

Using R as a converter to SPSS and Stata formats

If your team works in SPSS or Stata and wants native files, read the CSV or XPT in R and write them out. haven::write_sav(vs, "vital_signs.sav") writes an SPSS file and haven::write_dta(vs, "vital_signs.dta") writes a Stata file, carrying across any variable labels you have set. This runs on your side after export, and it is repeatable, so the conversion can be part of the analysis script.

Test your read-in script on sample data

Build a form in the free sandbox, enter a few sample subjects and export. Run the snippets above before the first real participant is screened.

Start building your study free

Workflow

A reproducible path from EDC to analysis

  1. 1

    Review the build with the SAP

    Check every endpoint component is a coded field, units are fixed and "not done" is separate from missing. The cheapest fixes happen before the first participant.

  2. 2

    Export from the sandbox

    Enter sample data, export raw and decoded files with the data dictionary, and note the exact file layout.

  3. 3

    Write the import script once

    Use the snippets above in R, SAS, SPSS, Stata or Python, and store it with the analysis code.

  4. 4

    Re-run at each cut-off

    For interim snapshots, filter by date range, freeze or lock the data, export, and re-run the same script. Record the export log entry in your data management file.

  5. 5

    Keep blinded and unblinded outputs separate

    Blinded roles receive exports without treatment arm. Plan how any unblinded outputs, for example for a data monitoring committee, will be produced by people with the right role.

Choosing a file

Which export for which job

JobUseWhy
Modelling and tablesRaw wide CSV with data dictionaryCoded values, stable column names
Data listings and reviewDecoded wide CSV or ExcelLabels instead of codes, readable by a clinician
Site-level or interim analysisEither, with site or date-range filterSnapshot matches the cut-off
Regulatory submission datasetsSDTM XPT with Define-XMLCDISC domains, one record per observation
Native SPSS or Stata filesRead CSV or XPT in R, then write_sav or write_dtaNo native writer in the platform, one extra step

FAQ

Questions teams ask before they switch

Something not covered here? Ask us directly.

How do I export EDC data to R from Capture?

Export the wide CSV or Excel file with its data dictionary, then read it with readr::read_csv or readxl. For SDTM datasets, read the SAS XPT files with haven::read_xpt. Example code for each is on this page.

Can Capture export directly to SAS, SPSS or Stata formats?

Capture exports CSV, Excel and SDTM SAS XPT files. It does not write native .sas7bdat or .sav files. SAS reads CSV with PROC IMPORT and XPT with the XPORT engine, SPSS reads CSV through the import dialog and XPT with GET SAS, and Stata uses import delimited or import sasxport5.

What is the difference between raw and decoded exports?

Raw files carry coded values, which is what you model. Decoded files carry labels, which is what reviewers read in listings. The data dictionary links the two.

How are not-done and missing values shown?

With markers: ND for not done, NA for not available, VND for visit not done and a separate not-applicable marker. An empty cell means the answer is missing.

Can I export only one site or one date range?

Yes. Exports can be filtered by date range and by site, which supports interim snapshots and site-level analyses.

Is every export recorded?

Yes. Each export goes into a separate append-only export log with user, role, time, filters, row count and whether it was blinded.

Do statisticians need a paid account to try this?

No. The sandbox is free with every feature, no credit card and no time limit. You pay only once you go live with real participants, so the export can be tested end to end beforehand.

Start building your study free

Build a form, enter sample data and test your R, SAS, SPSS or Stata read-in on a real export. No credit card, no time limit.

Start building your study free