Iowa Gambling Task

The paradigm on which the somatic-marker-hypothesis was tested, and the source of nearly every empirical claim in Bechara & Damasio (2005). Described in full in Bechara et al. (2000b), which is not in raw/; the account here is drawn from the 2005 review, with the original samples and two early manipulations from Damasio (1996), and the critical literature from Dunn et al. (2006).

Read the limitations block in the frontmatter before the account below. This page was built from the Iowa sources, and describes the task as they present it — a paradigm whose ambiguity forces subjects onto a non-cognitive strategy. Dunn et al. (2006) reviewed every published use of the task to July 2005 and concluded it “is no longer sufficient to be a major source of evidence for the SMH.” Their case is that the ambiguity is not as complete as claimed, the anticipatory SCR admits at least three readings, and the deficit has at least five candidate explanations. None of that stops the IGT being a useful and generative measure of decision impairment — which Dunn et al. say plainly. It stops the IGT being evidence for a specific mechanism.

What the task is for

The design goal is ambiguity, in the behavioural-economics sense (Einhorn & Hogarth 1985; Ellsberg 1961) — as distinct from certainty and from risk. Subjects never acquire the probabilities of reward and punishment, even once they know perfectly well which decks are good and bad. The level of uncertainty “remains high throughout.”

This is why the task carries the argument. Under stated odds, a decision failure can always be attributed to bad arithmetic. Under ambiguity, arithmetic is unavailable to everyone, so what separates good from bad performers has to be something else — which is the space the somatic marker is proposed to fill.

The four periods

The task’s most important result comes from stopping the game every 10 cards and asking subjects to say what they know (Bechara et al. 1997). Four periods emerge:

periodsubject’s knowledgenormal anticipatory SCRnormal behaviour
pre-punishmentsampling; no punishment yetnoneprefers high-paying A/B
pre-hunchpunishment encountered, “no clue”substantial risea hint of shift away from bad decks
hunchriskier decks suspected, not certainsustainedshift more pronounced
conceptualknows which decks are good and badsustainedadvantageous

The anticipatory SCR rises in the pre-hunch period — before any conscious knowledge, and alongside the first behavioural shift. VM patients never reported a hunch, never developed anticipatory SCRs, and kept choosing A and B.

The dissociation that does the work

The result the hypothesis rests on is not that VM patients do badly. It is the crossing of knowledge and performance:

  • 30% of controls never reached the conceptual period — and still performed advantageously.
  • 50% of VM patients did reach the conceptual period — and still performed disadvantageously.

So explicit knowledge is neither necessary nor sufficient. VM patients “may ‘say’ the right thing, but they ‘do’ the wrong thing.” Bechara & Damasio extend the dissociation to addiction (knows the consequences, takes the drug anyway) and psychopathy — an extension asserted from surface similarity, with no data, and worth resisting on this evidence.

The lesion double dissociation

groupSCR to reward/punishmentanticipatory SCRdeck choice
normalyesyes — larger before risky A/Bavoids A/B
VM lesionyes (slightly lower)noprefers A/B
amygdala lesionnonoprefers A/B

This is the pattern primary-and-secondary-inducers predicts: amygdala damage kills primary induction and therefore starves secondary induction of stored patterns; VM damage kills only the secondary trigger. Amygdala patients “can no longer register how painful it feels when one loses money,” which “misleads” the VM cortex about how painful a loss should feel.

Manipulations reported in the 2005 review

Emotional induction (Fig. 8). Ten healthy volunteers performed the task after recalling a neutral event (mowing the lawn) and after recalling an emotional one (the death of a loved one), order counterbalanced. Emotional induction reduced selections from advantageous decks. This is the entire empirical basis for the paper’s claim that emotion unrelated to the task at hand is disruptive — n = 10, no statistics reported, a bar chart with large error bars that visibly overlap.

Pharmacology. Dopamine and serotonin manipulation produce different effects on covert vs overt decision-making (Bechara et al. 2001) — DA biasing covertly via striatum, 5-HT overtly via ACC/SMA. The citation is a Society for Neuroscience abstract.

Three things only the 2000 source records

Added with the Bechara, Damasio & Damasio (2000) ingest — the Iowa laboratory’s own mid-point review, and the most patient-focused of the three primary statements this wiki holds.

The deficit is stable across four testings. Six VM patients and five controls were retested at one month, 24 hours later, and again at six months (Fig. 3). Controls improved significantly across repetitions; the VM group did not, and their net score declined. The authors read this as the laboratory counterpart of the patients’ real-life inability to learn from previous mistakes.

Worth holding at the same discount as everything else on this page — n = 6 and n = 5, no statistics, an odd retest spacing (1 month, then 24 h, then 6 months) that the paper does not explain. But it does bear on one live objection: if the VM deficit were a reversal-learning failure that extended practice could overcome, four exposures might have shown it. They did not. That is not an answer to Fellows & Farah (2005a), whose manipulation removes the reversal requirement rather than repeating the task, but it is a datum the reversal account has to accommodate.

The high-risk-taker label. The limitations block above records, via Dunn et al., that poorly-performing normals with intact anticipatory SCRs were classified “post hoc” as high-risk takers. The primary source makes the charge partly unfair and partly worse.

Partly unfair: these are described as “normal adults who describe themselves as high-risk takers in real life” — a self-report descriptor that is in principle available before the task, not a category invented to absorb the result. And the paper reports a psychophysiological distinction, not just a label: in these individuals anticipatory SCRs to the bad decks are slightly lower than to the good decks — the bias inverted — where advantageous performers show the normal direction, and where VM patients show no anticipatory SCRs at all. The claimed contrast is between an inverted-or-overridden bias and no bias, which is a claim about data.

Partly worse: no instrument is named for the self-description, no n is given, no statistics, no criterion for “slightly lower,” and no statement of whether the descriptor was collected before or after performance. So the correct charge is that the distinction is unverifiable, not that it is post hoc — and an unverifiable distinction load-bearing enough to explain away a third of the control sample is not obviously the better position.

Ageing, and its shape. Denburg et al. (1999): adults over 64 perform poorly relative to ages 26–56 — but dichotomously, some performing very well and some very poorly, rather than the group shifting as a whole. Recorded on age-related-interoceptive-decline, where this wiki’s aging thread reaches the same structure from the interoceptive side.

Two things only the 1996 source records

The reversed-design control (Anderson et al. 1996). Damasio (1996) names three candidate accounts of why VM patients take high immediate reward with severe delayed punishment — hypersensitivity to reward, insensitivity to punishment, or general insensitivity to future consequences — and reports a task variant built to separate them: punishment placed up front, with unpredictable reward schedules as the unexpected variable. The patients behaved the same way, which cuts against the first two accounts and leaves the third. This is the most useful methodological control in the Iowa programme and the 2005 review does not mention it. (Citation is a Society for Neuroscience abstract, in press at the time.)

Damasio also considers and rejects defective response inhibition as the sole source of the deficit — not on data, but on the grounds that the task’s complexity, the knowledge available to players, and the length of time over which the result holds make it implausible as a complete account. He grants there is much in animal and human work supporting the inhibition idea in general.

The conditioning dissociation. In the transcribed discussion, D. Bishop presses exactly the right objection: subjects do not merely learn a mapping, they must change one, since everyone starts on A and B because those decks pay more — so is this a failure to inhibit an earlier association rather than a failure to form associations? Damasio concedes the inhibition reading is conceivable, then reports that most patients who fail the gambling task acquire classical conditioning normally. If that holds, the gambling deficit is not a conditioning deficit and “conditioning” is not one thing. Filed also on pavlovian-defense-conditioning.

His actual reply to Bishop is the more revealing part: the intriguing thing is not whether inhibition failed but that such a failure is not compensated by the patients’ own realization that their strategy is losing — reasoning unaided by the marker does not prevail in guiding behaviour. That is the 1997 knowledge/performance dissociation stated as clinical intuition a year before Bechara et al. tested it.

Bishop was right, and someone eventually ran the experiment. Dunn et al. (2006) report Fellows & Farah (2005a), who did to the task exactly what Bishop’s objection implies: they rearranged the opening trials so the two disadvantageous decks no longer carried an early advantage, removing the reversal requirement while leaving everything else intact. VMPFC patients then performed the same as controls.

That is the strongest single result in the critical literature, and it lands on a specific claim. Damasio dismissed the inhibition account in 1996 not on data but on plausibility — the task’s complexity, the knowledge available to players, the duration over which the deficit holds. Nine years later the discriminating manipulation was run and the deficit disappeared when the reversal did. The conditioning dissociation above still stands (these patients do condition normally), so the honest statement is that “reversal” and “classical conditioning” are not the same capacity and the gambling deficit tracks the first.

Bechara et al. (2005) reply that the task contains more than contingency-reversal, and Turnbull et al. (in press, 2005) offer partial support — schizophrenic patients with negative symptoms acquired the standard task but failed added shift phases, suggesting the two are dissociable. Dunn et al. also note the framework can absorb reversal by claiming it needs an emotion-based “stop signal” — and then close that door: reversal learning survives amygdala lesions (Izquierdo et al. 2004), so it does not depend on emotion-processing structures.

Its replacement, built by its critics

Dunn et al. (2006) catalogued this task’s design flaws more thoroughly than anyone. Four years later Dunn et al. (2010) built the intuitive-reasoning-task — four decks, 100 trials, real money — that repairs them one by one: the reversal confound removed, deck position and card backs counterbalanced, magnitude crossed with profitability in a 2 × 2, the deck-exhaustion limit lifted, and the pre-decision window separated from the pre-feedback window.

Two results from that task bear directly on this page.

The first is a vindication. Crossing magnitude with profitability tests Tomb et al.’s (2002) variance interpretation — the objection in the limitations above that the anticipatory response may index reward variance rather than long-run goodness, since in this task the bad decks are also the high-variance ones. In the IRT they are not, and anticipatory bodily responses followed profitability with no magnitude effect and no interaction (Fs < 1). The body tracks goodness. That is the best available answer to Tomb et al., and it arrives from the critic’s side.

The second is a caution. 27% of a healthy sample still preferred the unprofitable decks (M = 25.74, SD = 40.71) — the same order as the 20–37% of IGT controls who perform disadvantageously, which this page flags as an ecological-validity problem. Fixing every design flaw did not fix that. Either a large minority of healthy adults really do decide badly under ambiguity, or four-deck card games measure something other than what their names claim, and no source in this wiki distinguishes the two.

What the IRT does not do is rescue this task’s data. Its clean answer to Tomb et al. is an answer about the IRT. In the IGT the magnitude/profitability confound remains exactly as Tomb et al. described it, and every published IGT result is downstream of it.

What the task cannot tell you

Two limits worth holding, both structural rather than fixable:

  1. It measures that a somatic state occurred, not which one. SCR is one sympathetic channel. Every claim in the framework about positive versus negative somatic states — including the whole background-somatic-states signal-to-noise model — is therefore unsupported by IGT data and imported from elsewhere. See autonomic-specificity-of-emotion.
  2. It rarely touches interoception, and never as the wiki needs. The task shows a bodily signal precedes and predicts good choices. It does not show that anyone perceives it. Pairing a decision task with the heartbeat-detection-task is what Dunn et al. (2010) did — and the answer matters: interoceptive accuracy has no relationship to decision quality (r = .08) but reverses the sign of the bodily signal’s effect depending on whether the body is right. A moderator that large, left unmeasured, sits in the error term of every result on this page. Werner et al. (2009) did pair this task with heartbeat counting and found good perceivers choose better — but as an extreme-groups main effect with the marker unmeasured, which is the Pollatos-style design the wiki discounts, not the moderation. So because Dunn measured the interaction against the intuitive-reasoning-task instead, and Werner measured only a main effect on the IGT, the moderator sits in this task’s error term still.