The large, placebo-controlled trial is no longer the default for digital therapeutics. As DTx evidence generation enters the synthetic control arm era, study teams are replacing sham applications with external control arms constructed from existing real-world data. But “synthetic control arm” can mean anything from a quick propensity score match to a fully pre-specified, regulatory-grade synthetic control arm from real-world data that satisfies regulators. This article walks through the practical decisions that separate exploratory analyses from evidence regulators can actually use.
Why the placebo arm is dying in DTx evidence generation
Sham-controlled designs were never a great fit for software-as-a-medicine. A sham app either looks different enough to break blinding or so similar that participants quickly realize they are in the control group. Even when blinding is technically possible, several problems persist:
- Engagement collapse: participants in a sham arm have little motivation to open the app, so the control experience no longer mirrors natural behavior.
- Ethical friction: delaying an effective digital therapeutic with a known mechanism to manage chronic conditions is harder to justify when data from prior versions and real-world users already exist.
- Software version drift: a DTx product changes over time. A sham control created at trial start may not represent the product version that will reach the market.
By 2026, regulators have made it clear: external controls are acceptable when a placebo is infeasible or unethical—provided the control arm is built with the same rigor as a randomized arm. The question is no longer whether synthetic controls can support DTx submission. It is how to make them regulatory-grade.
What does “regulatory-grade” mean for synthetic controls?
Regulatory-grade is not a certification. It is a property of the fit among the research question, the data source, the causal assumptions, and the transparency of the analysis. A synthetic control arm rises to that level when it:
- follows a protocol and statistical analysis plan written before outcome data are extracted;
- targets the same estimand as the intended claim or label;
- uses patient-level real-world data with clear provenance, variable definitions, and governance;
- addresses confounding explicitly through causal models, not just covariate adjustment;
- includes sensitivity analyses for missing data and unmeasured confounding;
- can be reproduced by an independent reviewer from raw data to final effect estimate.
In practice, this means treating the synthetic control arm like a trial component, not like a statistical afterthought.
Build from a target trial, not from a data source
Most weak synthetic controls start with a convenient dataset. Strong ones start with a written protocol for the hypothetical randomized trial that would be run if a sham were possible. This is the target trial emulation approach, and it is now the foundation of credible external control submissions.
Write the target trial protocol first
Specify the eligibility criteria, treatment arms, randomization ratio, endpoint definition, follow-up schedule, and analysis population before looking at real-world data. Every element of the target trial maps directly to how you will select and construct the synthetic control arm. For example, if the target trial excludes patients with severe comorbid psychiatric conditions, the real-world control cohort must apply the same exclusion rules with the same operational definitions.
Define the estimand, not just the endpoint
The estimand describes what is being estimated: the average treatment effect in the target population, the effect under a specific treatment regimen, or perhaps the effect of initiating treatment regardless of adherence. ICH E9(R1) estimand language is now expected in DTx submissions, and it matters even more for synthetic controls because you are not randomly assigning treatment. If your estimand is based on the initial treatment strategy, you need to handle intercurrent events—switching to another app, starting additional therapy, or upgrading to a new software version—through explicit strategies, not by deleting patients.
Choose real-world data with the right depth
Not all real-world data sources are suitable for building a regulatory-grade synthetic control arm. For digital therapeutics, the ideal source has patient-level longitudinal records, clinical and behavioral variables, and outcome ascertainment independent of the treatment assignment. Specifically, look for:
- timestamps and device-level usage logs that can be mapped to the DTx exposure window;
- baseline variables capturing symptom severity, duration of condition, prior treatments, comorbidities, and care setting;
- outcome measurements at clinically meaningful timepoints, not just at irregular visits;
- access to source documentation and an audit trail;
- sufficient overlap with the treated patient population after propensity score weighting or matching.
Claims data are useful for healthcare utilization and comorbid conditions, but they often lack the symptom scores and patient-reported outcomes that DTx labels require. Electronic health records can provide clinical details, but their visit frequency may not align with your endpoint window. Previous trial control arms and registry data can be excellent, provided you have patient-level records and the original informed consent allows reuse. One common mistake is relying on published aggregate summaries instead of raw patient-level data; synthetic control arms almost always require individual patient data for proper causal modeling.
Handle confounding, missingness, and non-adherence
The central challenge of a synthetic control arm is that treatment assignment is not randomized. Confounding must be addressed through design and analysis, and validation must be built into the study.
Create a credible propensity score model
Include variables that influence both the likelihood of receiving the DTx and the outcome. For digital therapeutics, that includes age, baseline symptom severity, educational level, digital literacy, prior treatment history, and comorbidity burden. Use propensity score weighting, matching, or stratification, and then check for sufficient overlap. If a large fraction of treated patients have no comparable real-world control, you are no longer estimating the intended effect.
Plan for informative missingness
Real-world data are rarely missing at random. A patient who stops filling out PHQ-9 forms may be worsening or recovering. A claim record without a follow-up visit may reflect disengagement, not stability. Use multiple imputation for some covariates, but handle outcome missingness with pattern-mixture models under several plausible assumptions about the missing data mechanism. A single complete-case analysis will not survive regulatory scrutiny.
Emulate intention-to-treat
In the target trial, all participants are followed regardless of adherence. In real-world data, patients may stop using the product or disappear from the dataset. To emulate intention-to-treat, you need to keep those patients in the control arm and use time-varying weights or censoring methods for treatment switching. Removing non-adherent patients from the synthetic arm re-introduces selection bias and creates an unfair comparison against the DTx arm where non-adherent participants are still counted.
Pre-specify the analysis and challenge the synthetic control
The most persuasive synthetic control arm is pre-specified in detail before any outcome data are extracted. The analysis plan should include the covariate adjustment model, matching or weighting method, caliper width, outcome model, handling of ties, and planned sensitivity analyses. It should also include negative control outcomes: events that the DTx should not plausibly affect, such as an unrelated injury diagnosis or a routine administrative claim. If the synthetic control shows an unexpected effect on a negative control, there is likely residual confounding.
Another useful validity check is comparing the synthetic control arm’s outcomes to those from an internal run-in group or to published reference estimates for the same population. Triangulating with a second real-world source—such as a different claims database—adds credibility. Regulators are less interested in a single narrow p-value than in the full picture of uncertainty across model specifications and assumptions.
Conclusion
The synthetic control arm era is not about replacing randomized evidence with convenience data. It is about applying randomized trial discipline to external data. When designed with a target trial protocol, patient-level real-world data, rigorous causal methods, and pre-specified diagnostics, a synthetic control can provide the reliable comparative evidence DTx developers need—without asking a single participant to stare at a sham app.
