The most widely cited number in the caregiving literature, a 63 percent rise in mortality, belonged to strained elderly spouses rather than to caregivers as a group; a propensity-matched national cohort later found caregivers outliving their matches. The record also shows what the wound-healing and telomere studies actually measured, and which supports passed randomized trials.

The house is dark except for a strip of light under one door. Someone is awake at ten minutes to five, and not by choice: she is listening for a cough from the next room, counting backward to the last pill, running tomorrow’s arithmetic in her head, the pharmacy that closes at six, the shower that now takes two people, the breakfast that has begun to require twenty minutes of coaxing. Nobody assigned her this shift. It assembled itself around her, one small emergency at a time, until it had a schedule and she had a second job with no wage, no clock-out, and no title beyond a word she rarely uses for herself.
Epidemiology has a flat name for the person behind that door — the informal caregiver — and she is a composite rather than an individual, an archetype standing in for the millions of households that run on her unpaid hours. For a quarter century the research literature has been arguing with itself about what the role does to the body performing it, and the argument has a number at its center: 63 percent. It may be the most quoted statistic in the field, and it is routinely quoted wrong.
Two stories circulate about what caregiving does to a body, and they cannot both be the whole truth. In one, caregiving is a slow-motion injury: it depresses immune function, ages cells ahead of schedule, and raises the risk of death by nearly two-thirds. In the other, caregivers as a group actually outlive comparable people who never take on the role. Both stories cite peer-reviewed studies, and both are told by careful researchers. The work of this article is to lay the two side by side, the strain biology and the population arithmetic, and to mark where each is strong, where each is thin, and what the randomized trials say about the only question that is fully practical: what helps.
The number comes from the Caregiver Health Effects Study, published by Richard Schulz and S. R. Beach in JAMA in December 1999. The design was patient and unglamorous: a prospective, population-based cohort in four American communities, followed from 1993 through 1998, on average about four and a half years. The participants were 392 older adults living with and caring for a spouse, and 427 noncaregivers, all between 66 and 96. What made the study durable was a distinction most summaries of it discard. Schulz and Beach did not treat caregiving as one exposure; they sorted the households into four situations: spouse not disabled; spouse disabled but receiving no help from the participant; spouse disabled, helped, with the helper reporting no strain; and spouse disabled, helped, with the helper reporting mental or emotional strain.
That four-way split is the study’s real contribution, because it separates three things that usually travel together and get blamed as one: living with illness, doing the physical work of care, and experiencing the work as strain. A husband can spend his mornings helping his wife bathe and dress and report, honestly, that he is managing; his neighbor can perform the identical task list and feel himself coming apart. On paper the two men are the same exposure. Schulz and Beach bet that the difference between them, the internal weather rather than the chore chart, was where the health effect lived, and they built the cohort so the bet could be tested.
After four years, 103 participants, or 12.6 percent, had died. Once the investigators adjusted for demographics, existing disease and subclinical cardiovascular disease, one group stood apart: caregivers who reported strain had a mortality risk 63 percent higher than the noncaregiving controls. The relative risk was 1.63, with a 95 percent confidence interval of 1.00 to 2.65, an interval whose lower edge rests exactly on the line of no effect. The other groups told the quieter half of the story. Caregivers doing the same work without strain came in at a relative risk of 1.08, statistically indistinguishable from people not caregiving at all, and spouses of disabled partners who were not providing care landed at 1.37, also not significant.
Closely read, the study said something narrower and more interesting than its reputation: among elderly spousal caregivers, what predicted death was not the caregiving but the strain, and the unstrained, performing the same tasks in the same kind of household, looked like everyone else. That distinction did not survive contact with the wider world. A 2015 review in The Gerontologist notes that the finding became widely cited as evidence of the physical health risks of caregiving and a centerpiece of advocacy for caregiver services — shorthand, in effect, for the claim that caregiving, itself, kills.
The biological case that chronic stress leaves marks on the body is real, and it rests on some of the strangest small studies in the stress literature. The most vivid comes from Ohio State University, published in The Lancet in 1995. J. K. Kiecolt-Glaser and colleagues recruited thirteen women caring for relatives with Alzheimer’s disease, their average age about 62, and thirteen women matched for age and family income who were not caregiving. Every woman received the same standardized injury, a 3.5-millimeter punch-biopsy wound, one identical small circle of missing skin. Then the researchers photographed the wounds and waited.
The waiting had a protocol. Healing was tracked by photography and by a bubbling test: hydrogen peroxide foams on broken tissue, and the study defined a wound as healed when the peroxide no longer foamed. It is the kind of operational definition that makes the study feel less like psychology and more like carpentry, with no questionnaires and no self-report, only skin that either had or had not closed.
The controls healed in 39.3 days on average; the caregivers took 48.7, roughly nine days longer, about a quarter again as long, for the same wound in matched women. Their blood told a parallel story: stimulated in the lab, the caregivers’ white cells produced significantly less messenger RNA for interleukin-1β, a signaling molecule involved early in tissue repair. Twenty-six women is a very small study, and it deserves to be held lightly. But the image has endured in the field for a simple reason: the wound was the same, and the healing was not.
A second landmark works at the level of the cell, specifically at the ends of chromosomes, where repeating DNA caps called telomeres shorten a little with each cell division until the cell retires. Telomere length is one crude clock of biological aging, and telomerase is the enzyme that maintains it. In 2004, Elissa Epel, Elizabeth Blackburn and colleagues reported in PNAS that among healthy premenopausal women, psychological stress — both its perceived intensity and how long it had gone on — was significantly associated with higher oxidative stress, lower telomerase activity and shorter telomeres in blood immune cells. Women with the highest perceived stress carried telomeres shorter, on average, by the equivalent of at least a decade of additional aging compared with the least stressed. It is an association study, a mechanism sketch rather than a verdict, but it gave the stress hypothesis its cellular vocabulary.
A meta-analysis widens the view, and in it the proportions come into focus. In 2003, Martin Pinquart and Silvia Sörensen integrated 84 articles comparing caregivers with noncaregivers. The psychological differences were moderate and consistent: effect sizes of 0.58 for depression and 0.55 for stress, with subjective well-being lower by 0.40. The difference in physical health, though, was small, at 0.18: statistically significant but modest. Across outcomes, the caregiver–noncaregiver gaps ran larger in dementia-caregiver samples than in mixed ones. The robust, repeatable finding is the psychological toll; the physical-health signal, averaged across everyone who carries the label, is far thinner than the folklore.
Before the counter-evidence comes a word about how studies find caregivers in the first place, because the method decides what a study can see. A sample recruited through memory clinics, support groups and advocacy networks skews toward the people those places exist for, the ones deep in dementia care and already strained enough to seek help. A study that begins instead from a population cohort, thousands of people enrolled for other reasons, some of whom happen to be caregivers, takes in the full range, including the majority managing without crisis. Neither lens is wrong; one is a close-up of the hardest cases, the other the wide shot. Some of the apparent contradiction in this literature is two lenses being mistaken for one, though not all of it: the two mortality studies at the center of this story were both population-based, which is exactly what makes their disagreement worth taking seriously.
The literature then did what literatures are supposed to do and rarely get credit for: it checked itself at scale. The Reasons for Geographic and Racial Differences in Stroke (REGARDS) study, a national cohort, contained 3,503 family caregivers. In an analysis published in the American Journal of Epidemiology in 2013, David Roth and colleagues matched each one to a noncaregiver with the same propensity profile across fifteen demographic, health history and health behavior covariates, then followed both groups for about six years.
The caregivers died less. Two hundred sixty-four caregivers, or 7.5 percent, died during follow-up, against 315, or 9.0 percent, of their matches; as a hazard ratio, 0.823, with a confidence interval of 0.699 to 0.969 — an 18 percent lower rate of death. The subgroup analyses are the part worth underlining: cuts by race, sex, caregiving relationship and reported caregiving strain “failed to identify any subgroups with increased rates of death” compared with matched noncaregivers. Even the strained caregivers, in this larger and later sample, were not dying faster than their matches.
Propensity matching deserves a sentence, because the result leans on it. Rather than comparing caregivers to whoever happened not to be one, the analysis built each caregiver a statistical twin, a noncaregiver with the same profile across fifteen measured dimensions of demographics, health history and health behavior, and raced the pairs forward through time. Matching can only balance what is measured, and one candidate explanation it cannot fully retire is selection: a person usually has to be reasonably well to take on the work at all, so some of the survival edge may belong to who becomes a caregiver rather than to what caregiving confers. That is a real limit. It is also beside the central point, which is negative: even under this design, no subgroup could be found dying faster.
The two flagship studies genuinely disagree about strained caregivers, and the honest reading is not that one refutes the other. Schulz followed elderly spouses in the 1990s; Roth matched family caregivers of all kinds in a national cohort a decade later; the eras, the designs and the samples all differ. What the pair establishes together is the boundary of the claim: elevated mortality is not a general property of the caregiving role, and where it has been observed at all, it has been in the oldest, most strained spousal corner of the picture — the corner Schulz himself flagged, in a finding whose confidence interval touched the null.
By 2015, Roth, Lisa Fredman and William Haley could count five population-based studies since Schulz and Beach finding reduced mortality and extended longevity for caregivers as a whole, and they noted that most caregivers report benefits from the role and that many report little or no strain. Their review argued that policy reports and media coverage present “an overly dire picture” while largely ignoring the positive findings, and that the practical move is not reassurance for its own sake but triage, aiming the evidence-based services at the subgroup that is highly strained. The correction is a decade old, and outside the journals almost nobody has heard it.
If strain is the exposure that matters, the operative question becomes whether anything reliably reduces it. Here the record is unusually concrete, because caregiver support has been tested the way drugs are tested, in randomized controlled trials, and the results split cleanly by the kind of support offered.
The largest is REACH II, published in the Annals of Internal Medicine in 2006: 642 in-home dementia caregivers in five American cities, recruited in three roughly equal groups — 212 Hispanic or Latino, 219 white, 211 Black or African-American — and randomized to a structured program or to two brief check-in calls. The program was work: twelve in-home and telephone sessions over six months, aimed at depression, burden, self-care, social support and the care recipient’s problem behaviors. At six months, the quality-of-life composite built from those same domains had improved significantly more in the intervention group among Hispanic or Latino caregivers, white caregivers, and Black spouse caregivers. The starkest single result was clinical depression, at a prevalence of 12.6 percent in the intervention group against 22.7 percent in the controls. What the program showed no statistically significant effect on, at six months, was the rate of nursing-home placement.
The trial’s own fine print is worth carrying along: a single follow-up assessment at six months; cultures and ethnicities grouped into three broad categories; some groups not enrolled at all; and, in the Black or African-American arm, a significant quality-of-life gain specifically among spouse caregivers. Six months is long enough to show that a structured program moves depression and daily quality of life; it says nothing about years, and the investigators did not pretend otherwise.
A different trial moved that last outcome too. At New York University, Mary Mittelman’s group enrolled 406 spouse caregivers of community-dwelling Alzheimer’s patients over nine and a half years and randomized them to usual care or to an enhanced program: six sessions of individual and family counseling, participation in a support group, and a telephone line to counselors that stayed open indefinitely, help available at the moment of crisis rather than by appointment. Published in Neurology in 2006, the result was a 28.3 percent reduction in the rate of nursing-home placement (hazard ratio 0.717), with the model-predicted median time to placement arriving 557 days later — about a year and a half. The mediation analysis carried the meaning of the result: improvements in the caregivers’ satisfaction with social support, their response to the patient’s behavior problems, and their own depression symptoms together accounted for 61.2 percent of the effect. Supporting the caregiver was itself the mechanism.
Then there is respite, the scheduled break from caregiving, the support most commonly advocated and most intuitively obvious. A 2014 Cochrane review found four randomized trials with 753 participants, too different from one another to pool, with evidence rated very low quality overall, and no significant effect of respite versus no respite on any caregiver variable. The reviewers were careful about what that means: it “may reflect the lack of high quality research in this area rather than an actual lack of benefit.” The most recommended support is, for now, the least well-supported by high-quality evidence, an honest null worth knowing before anyone treats a few hours off as the whole answer.
Laid side by side, the three results form a pattern. The interventions that passed their trials were not breaks from the role; they were apprenticeships in it: twelve sessions on depression, burden, self-care, social support and difficult behaviors; counseling that pulled the whole family into the room; a phone line with no closing time. Each treats caregiving as skilled work that can be taught, supported and shared, and treats the caregiver’s own mind as a legitimate clinical target. The intervention that treats caregiving as a burden to be paused, worthy as it sounds, is the one that has not yet demonstrated an effect. The distinction is not a verdict on respite; it is a map of where the evidence currently is.
| Support tested | Trial / review | What it showed |
|---|---|---|
| Multicomponent skills program — 12 in-home/phone sessions over 6 months | REACH II RCT, 642 dementia caregivers (2006) | QoL composite improved among Hispanic/Latino and white caregivers and among Black spouse caregivers; clinical depression prevalence 12.6% vs 22.7%; no significant difference in placement at 6 months |
| Counseling + support group + always-available phone counseling | NYU RCT, 406 spouse caregivers (2006) | Placement rate reduced 28.3% (HR 0.717); model-predicted median placement delayed 557 days; 61.2% of the effect mediated by support, coping and mood improvements |
| Respite care (scheduled breaks) | Cochrane review, 4 RCTs, 753 participants (2014) | No significant effect on any caregiver variable; evidence very low quality — a research gap more than a proven null |
Held in one frame, the twenty-five years describe a picture more precise, and more humane, than the folklore in either direction. The role itself is not associated with shorter life in population samples; in the largest matched analysis, caregivers outlived their matches by a modest margin. The exposure to watch is strain — the specific, self-reported state Schulz singled out — while the laboratory’s adjacent findings, slower wound healing in a small sample of caregivers and shorter telomeres in women under high chronic stress, show what sustained psychological stress in general can write into tissue. The supports with the strongest randomized evidence are not vague encouragements to self-care but structured programs: skills sessions in the home, counseling that includes the family, a phone that answers at midnight. In the NYU trial, most of the measured benefit for the patient ran through the caregiver’s own support, coping and mood.
What is still missing is worth naming plainly. Strain, the pivotal variable, is measured in these studies by asking; Schulz’s cohort simply reported whether the helping came with mental or emotional strain, and a science that turns on self-report will always have soft edges. The strongest support trials stop at six months or track a single outcome, and the most advocated support has four trials, totaling 753 participants, to its name. The 2015 reappraisal’s prescription follows from all of it: a public picture more balanced than the dire one, and evidence-based services aimed where the risk actually concentrates, at the caregivers who are highly strained or otherwise at risk.
The story returns, in the end, to the strip of light under the door. The person standing behind it is statistically likely to be doing better than the headlines about her suggest, and if she is not, the literature’s clearest finding is that her depression, her coping and her day-to-day quality of life are the tractable part: the variables randomized trials have actually moved, and, in the NYU analysis, the channel through which most of the measured benefit ran. In the strongest studies, support for the caregiver was not a kindness added onto the care plan; it was the intervention itself.
Educational, not medical advice.
The famous 63 percent mortality increase (Schulz & Beach, 1999) applied only to elderly spousal caregivers reporting strain — unstrained caregivers showed no elevation (RR 1.08) — and a propensity-matched national analysis of 3,503 caregivers later found an 18 percent LOWER death rate than matched noncaregivers, with no at-risk subgroup even by strain. The reliable toll is psychological (meta-analytic effect sizes ~0.55–0.58 for stress and depression) with a small average physical-health difference (0.18); mechanism studies link chronic psychological stress to slower wound healing in caregivers (48.7 vs 39.3 days for an identical biopsy wound) and to shorter telomeres in high-stress women. Randomized trials back structured support: REACH II's 12-session program cut clinical depression prevalence from 22.7% to 12.6%, and NYU's counseling-plus-support-group program reduced the nursing-home placement rate 28.3%, delaying the model-predicted median time to placement by 557 days, with 61.2 percent of the effect running through the caregiver's own support, coping and mood. Respite care, the most-advocated support, remains an evidence gap (Cochrane: 4 trials, 753 participants, very low quality, no demonstrated effect).
9 peer-reviewed sources, published 1995–2015, across 9 journals. Every citation links to its PubMed record.
Each links to its Magellan monograph — what it is, what it does, and the studies behind it.
An exercise-linked signal produced intriguing results in Alzheimer’s mouse models. Here is what the…
Rentosertib has renewed the debate over biological age. A randomized human experiment offers an important…
Tea, berries and spices were associated with a changing DNA-methylation clock in a small lifestyle trial.…
Prefer the interactive version? Open this article inside the Magellan app →