Seventeen clinician-rated items, some scored 0 to 4 and some 0 to 2, for a total of 0 to 52. The HAM-D has been a standard efficacy endpoint in depression trials for decades. Record each item so the total and response can be verified.
Free sandbox · No credit card · 21 CFR Part 11 aligned
Week 8 total
9 / 52
At a glance
In trials
Depression trials use two main clinician-rated scales: the HAM-D and the Montgomery–Åsberg Depression Rating Scale (MADRS). The HAM-D includes more somatic items such as sleep, appetite and anxiety symptoms, which some argue makes it sensitive to side effects as well as mood; the MADRS focuses more on core mood symptoms. Many recent trials use the MADRS as primary and the HAM-D as secondary, or the reverse, depending on the development programme.
Whichever is chosen, a self-report measure such as the PHQ-9 is often collected alongside, together with the CGI as a global rating and the C-SSRS for suicidality.
The original HAM-D gives item anchors but no interview script, which leads to rater variability. Structured interview guides were developed to standardise questioning. Specify the guide in the protocol and train all raters on it.
1. Depressed mood (0 to 4)
HAMD014. Insomnia: early (0 to 2)
HAMD0417. Insight (0 to 2)
HAMD17Structure
| Items | Content | Range |
|---|---|---|
| 1 to 3 | Depressed mood, guilt, suicide | 0 to 4 each |
| 4 to 6 | Insomnia: early, middle, late | 0 to 2 each |
| 7 to 8 | Work and activities, retardation | 0 to 4 each |
| 9 to 11 | Agitation, psychic anxiety, somatic anxiety | 0 to 4 each |
| 12 to 14 | Gastrointestinal and general somatic symptoms, genital symptoms | 0 to 2 each |
| 15 | Hypochondriasis | 0 to 4 |
| 16 to 17 | Loss of weight, insight | 0 to 2 each |
Severity bands for the total vary between publications. Pre-specify the ones you use.
In Capture
Each item has only its permitted values, so a 0 to 2 item cannot be given a 3. The rater is recorded on the form, and the field-level audit trail records who entered each value and any later change with a reason.
Edit checks / auto-queries
2Type
Range High
Operator
Greater than
Value
180
Priority: High
Type
Range Low
Operator
Less than
Value
80
Priority: Normal
Query raised automatically
Value 192 violates limit (180). Please verify.
Copy the template into the free sandbox.
Signal detection
High placebo response is one of the main reasons depression trials fail. Rater behaviour contributes: raters who know a participant must score above a threshold to enrol may unconsciously inflate baseline scores, which then fall at the next visit regardless of treatment. Common safeguards include independent raters at baseline, central or remote rating, and review of score patterns across sites.
In the data, that means knowing who rated each assessment, when, and whether the rating was done remotely. It also means looking at the distribution of baseline scores just above the eligibility threshold. See risk-based monitoring and SDV for central data review.
Before go-live
17-item or longer version, with the scoring stated.
Structured guide named in the protocol; terms checked.
Certification before rating participants.
0 to 4 and 0 to 2 items enforced.
Definitions in the analysis plan.
Independent or central rating where planned.
Usually about 15 to 20 minutes with a structured interview guide.
Usually the past week, as defined by the interview guide in use.
Nine items are scored 0 to 4 and eight items 0 to 2. The sum gives a total of 0 to 52.
A total of 7 or less is the most common remission definition.
Usually as a reduction of 50% or more from the baseline total.
The original 1960 scale is in the public domain. Structured interview guides may have their own terms.
Trained clinicians. Most trials require rater certification.
Both are widely used. The HAM-D includes more somatic items; the MADRS focuses more on core mood symptoms.
Per-item ranges, rater attribution, audit trail. Free sandbox.