Between the 1980s and the early 2000s, John Gottman and Robert Levenson ran a series of observational studies of married couples, first at the University of Illinois and later in the Seattle laboratory that press coverage came to call the Love Lab. The procedure was consistent: a couple discussed an area of ongoing disagreement for around fifteen minutes while the conversation was recorded and physiological measures were taken. The recording was then coded moment by moment, using the Specific Affect Coding System, which assigns each segment of the exchange to a category of emotional behaviour.1

Two quite different things came out of that programme. They have very different evidential standing, and popular coverage has tended to report them as a single result.

The behavioural categories

The first output is descriptive. The coding surfaced a small set of recurring behaviours in conflict conversations — criticism, contempt, defensiveness and withdrawal — which Gottman later grouped under a memorable collective label in his books for general readers.

As descriptive categories these have held up reasonably well. Independent groups using different coding schemes have identified comparable behaviours in recorded disagreements. The demand–withdraw pattern in particular, in which one partner presses an issue while the other disengages, was documented by Andrew Christensen and Christopher Heavey working separately from the Gottman group, and has since been observed across many samples.2 Of the four, contempt is the category most consistently associated with poorer later outcomes.

That association is a pattern observed across samples, not a mechanism established by experiment. Couples were not assigned to interaction styles; they arrived with them.

The accuracy figures

The second output was predictive, and it is where the difficulty sits. Figures above ninety per cent — the share of couples whose later separation a statistical model correctly identified — travelled widely through press coverage and popular writing.

Methodologists have objected for two decades that those figures were generally not produced the way the word prediction implies. In most of the published analyses, the model was fitted to the same sample whose outcomes it then described. A model with enough parameters, specified after the outcomes are known, will describe almost any dataset closely. Prediction in the strict sense requires the model to be fixed in advance and then tested against a sample it has never seen.

Richard Heyman and Amy Smith Slep set this out directly in the methodological literature, distinguishing post hoc classification from prospective prediction and showing what happens to the figures when cross-validation is applied.3 The journalist Laurie Abraham revisited the same problem for a general readership in 2010, tracing how the numbers moved from journals into headlines with the qualifications removed along the way.4

Where independent groups have attempted something closer to a prospective test, the results have been weaker. Hyoun Kim, Deborah Capaldi and Lee Crosby applied the approach to a separate longitudinal sample and reported predictive performance well below the headline figures.5

Why the distinction matters

A descriptive finding — that couples who later separated had shown more contempt in a recorded conversation — is entirely compatible with the accuracy claim being overstated. Both statements can hold at once. Coverage has tended to collapse them, so that a modest observational regularity arrives in public discussion carrying a number it was never entitled to.

What the design cannot establish

Beyond the statistical objection, the design itself limits what can be concluded. Any behavioural difference observed in the laboratory might be a cause of later outcomes, a consequence of conditions already present, or a visible marker of something else entirely — chronic financial strain, untreated illness, caregiving load.

Sample composition is a further limit. The laboratory samples were drawn largely from volunteers in specific North American university towns and skew white, educated and middle class. Whether the same coded behaviours carry the same weight in other populations is a question these designs were not built to answer.

Sources

  1. Gottman, J. M., & Levenson, R. W. (1992). Marital processes predictive of later dissolution. Journal of Personality and Social Psychology. See also Gottman & Levenson (2000), Journal of Marriage and Family.
  2. Christensen, A., & Heavey, C. L. (1990). Gender and social structure in the demand/withdraw pattern of marital conflict. Journal of Personality and Social Psychology.
  3. Heyman, R. E., & Smith Slep, A. M. (2001). The hazards of predicting divorce without crossvalidation. Journal of Marriage and Family.
  4. Abraham, L. (2010). Can you really predict the success of a marriage in fifteen minutes? Slate.
  5. Kim, H. K., Capaldi, D. M., & Crosby, L. (2007). Generalisability of Gottman and colleagues' affective process models of couples' relationship outcomes. Journal of Marriage and Family.