Is interoception domain-general or organ-specific?
Created with the Banellis et al. (2026) ingest, and overdue in the same way is-the-heartbeat-counting-task-valid was overdue: the material had been accumulating for six ingests — Ferentzi on interoceptive-taxonomy, the “scarce and inconsistent” line on quadt-2018-interoception-health-disease, the exceptionality argument on respiratory-interoception, the pending-question note on interoceptive-control — without the fight being named.
Why it is the more damaging of the wiki’s two measurement debates
is-the-heartbeat-counting-task-valid asks whether one instrument works. This asks whether the construct travels. They are independent, and the second is worse for this wiki, because a bad instrument can be replaced while a non-existent general construct cannot.
Consider what the wiki routinely does. Dunn measures cardiac counting accuracy and concludes about intuitive decision-making. Wiens measures cardiac discrimination and concludes about emotional intensity. Nentjes measures cardiac discrimination and concludes about psychopathy. Khalsa measures cardiac discrimination and concludes that interoception declines with age. Each conclusion is stated about interoception; each measurement is about a heart. If the axes are independent, every one of those inferences needs its scope narrowed, and the is-more-interoceptive-awareness-better collision table — which lines up findings from different labs, disorders and outcomes as though they were measurements of one quantity — is comparing quantities that have been shown not to covary.
What is agreed
- First-order performance does not travel. Nobody in this debate now claims cardiac accuracy predicts respiratory accuracy. Even Garfinkel et al. (2016), the domain-generality anchor, reported first-order performance as unrelated across the axes; their claim was always about the metacognitive layer.
- Something does travel: confidence. Both camps find it. The dispute is over what it is.
- The instruments were incommensurable until recently. Comparing a counting score to a filter-detection score confounds modality with method. The HRDT/RRST pair is the first matched instrumentation, which is why the 2026 null carries more than the earlier ones.
The crux: what is the domain-general confidence signal?
This is where the debate now actually lives, and both readings fit the data.
| Confidence as interoceptive sensibility | Confidence as general metacognitive bias | |
|---|---|---|
| What the cardiac↔respiratory r = 0.51 shows | A real shared interoceptive layer — the person’s relationship to their body as a whole, above the organ-specific perceptual machinery | A response-scale disposition with no bodily content |
| What the respiratory↔auditory r = 0.64 shows | An embarrassment: the largest correlation involves a task with no body in it | Exactly what is predicted; audition is not special, nothing is |
| Consequence for maia and questionnaire measures | They measure the general interoceptive layer, and their failure to predict accuracy is expected | They measure a trait that is not interoceptive at all, and calling it interoceptive sensibility is a category error |
The auditory correlation being the largest is the strongest single fact in this debate and it favours the right-hand column. A shared interoceptive layer should not bind more tightly to tones than to breath.
The wiki’s caution against overreading it: mean confidence on any two VAS-rated forced-choice tasks will correlate partly through scale use, and no study here decomposes that. “Confidence is domain-general” and “confidence ratings share method variance” predict the same matrix.
The second leg: the brain does not treat the channels alike either (Haruki & Ogawa 2023)
Everything above is an argument from individual differences — scores from different people, correlated or not. That evidence class has a standing weakness this debate has not had to face: a null correlation can be produced by unreliable measurement rather than by independent faculties, and the instruments here are contested on exactly that ground (see is-the-heartbeat-counting-task-valid).
Haruki & Ogawa (2023) supply evidence of a different kind, and it is within-subject. The same 31 people attended to their heart and to their stomach minutes apart, against the same control. Reliability of a difference score across people is not at issue; what is measured is whether one brain does the same thing twice.
It does not:
- Right dorsal anterior insula was the only region preferring cardiac attention — and this wiki’s most-cited interoceptive region turns out to prefer one organ (see insular-cortex).
- Gastric attention recruited occipitotemporal visual cortex overlapping the gastric-network, plus orbitofrontal cortex, hippocampus and primary motor cortex — a feeding-and-foraging profile with no cardiac counterpart.
- Left dorsal middle insula classified which organ was being attended to from its multivoxel pattern, while showing no mean-activation preference for either.
Why the two legs are independent, and why that matters. Decorrelated individual differences are compatible with a shared substrate (two people could use the same circuit with different competence). A shared substrate is compatible with decorrelated scores (same machinery, channel-specific signal quality). Neither result entails the other; they are separate predictions of the organ-specific view, and both came out.
What it does not do is fill the gastric gap below. This is a neural signature, not a psychophysical axis. The three-channel test still requires a gastric task with a correct answer, and this paradigm has none — it measures attention, not accuracy (interoceptive-attention-task). The gastric channel now has a brain signature and still no way to be good or bad at it.
And one caution against overreading in the domain-specific direction. Two channels should differ somewhere in the brain — they carry different information from different organs for different purposes, which is true of the visual and auditory systems without anyone concluding that exteroception is not a thing. The finding that bites is not that the maps differ but that they differ at the top, in the anterior insula where the field locates integration and unified feeling. Divergence at the transduction end would be unremarkable; divergence at the re-representation end is what a domain-general account cannot easily absorb.
The third leg: a sign reversal, which a null cannot give you
The two legs above are (i) decorrelated individual differences and (ii) divergent neural substrate. Both are arguments from absence — no correlation, no shared map. Both are therefore vulnerable to the standing objection that unreliable instruments manufacture nulls.
Harrison et al. (2021) supply evidence of a third kind, and it is immune to that objection in a way the others are not.
| channel | instrument | anxiety and detection |
|---|---|---|
| cardiac | counting / discrimination | anxious and panic samples detect better |
| respiratory | FDT | anxious participants detect worse (Z = −2.4, p = 0.01) |
Noise attenuates correlations toward zero. It does not flip their sign. So if both findings are real, the channels are not merely independent — a single trait relates to them oppositely, which no version of “one interoceptive ability, differently measured” can absorb.
The honest caveat is that the two findings come from different instruments as well as different organs, so the confound this debate has fought about all along is still present. And the alternative reading is not friendly to the cardiac side: the counting task has a documented mechanism by which anxiety produces a higher accuracy score without better perception (count faster in a population that undercounts — see anxiety-sensitivity). The FDT has no such route; the load is physically present or absent. On that reading there is no reversal to explain, only one real finding and one artefact — which would be worse for the cardiac literature than for domain generality.
The discriminating study is small and has not been run: both tasks, one anxiety-selected sample. It is easier than the isoproterenol experiment ranked first below, and it would settle whether this page has a reversal or a broken instrument.
The fourth leg, pointing the other way: a positive correlation under perturbation
Everything above argues for organ-specificity. Wilzok et al. (2023) is the page’s one contrary datum and it should be held at its actual weight — which is less than it looks, and more than nothing.
Fifty-two people did the same two paradigms twice: once with graded pinprick pain, once with graded inspiratory loading. Same grey-shade cues, same three VAS, same trial structure, same session. The quantity was the anticipation–experience gap, manipulated by delivering a stimulus one level above or below what the cue promised. Across channels it correlated at r = 0.57.
Three things stop this from being the domain-generality result the debate has been waiting for.
It is at the wrong layer. The decorrelation findings are about sensitivity, precision and metacognitive efficiency — first-order perceptual competence and its readout. A discrepancy score is the difference between two marks a person makes on the same scale seconds apart. Note where the wiki’s other cross-modal positive sits: confidence, also a self-report magnitude, also not a threshold. The pattern across both studies is that self-report quantities travel across channels and psychophysical ones do not, which is a tidier summary of the evidence than either paper offers and is not friendly to a general ability.
Shared method variance is unexcluded and the design maximizes it. A participant with a consistent anchoring habit produces correlated discrepancies with no shared bodily machinery. This is the same caution the crux table applies to Banellis’s confidence result — but there it could be tested, because Banellis included an auditory task and the auditory correlation turned out to be the largest. Wilzok has no exteroceptive comparison. The decisive control is exactly the one missing.
One of the two channels may not qualify. See nociception. On Ceunen et al.’s narrow definition, cutaneous pinprick is somatosensory and this is not a within-interoception correlation at all. The authors adopt broad inclusion explicitly and the finding depends on it.
What survives all three brakes is that both channels were aversively perturbed — and that is the condition escape 1 says has never been tested. It is a consistency, not a test, since arousal is confounded with layer, instrument and channel set. But the escape ranked first below now has one observation pointing its way instead of none.
The escapes, ranked
- Arousal / perturbation (strongest, and now with one datum). All the null evidence is from resting healthy participants. Predictive accounts hold that interoceptive prediction errors matter most when the body is perturbed; a common central factor could exist and be invisible at rest. sahib-khalsa’s isoproterenol programme is built on this premise, and the Banellis authors endorse the test. This is still the discriminating experiment the debate needs and does not have — but see the fourth leg above: Wilzok et al. perturbed two channels hard and got r = 0.57, which is what this escape predicts and is not yet evidence for it, because nothing in that study varies arousal while holding the measure fixed. The experiment is now cheaper to specify than it was: one sample, one battery, rest versus challenge.
- Development and pathology. All the samples are healthy young adults. A general factor could be present in a clinical population, or in childhood before channel-specific expertise diverges — though Chen et al. note there is no instrument for non-verbal populations, so the developmental version may be untestable with current tools.
- Precision-weighting (Garfinkel’s). Channels may be differentially weighted rather than differentially able, so a person could be perceptually competent everywhere and behaviourally sensitive only where weight is assigned. This predicts decorrelated task performance with a shared underlying capacity — i.e. it is compatible with every null on this page while preserving a general mechanism. Elegant, and hard to falsify.
- Task asymmetry (weakest but unresolved). HRDT is multisensory, RRST unimodal. Fixable, and the authors say how.
What would resolve it
- The same battery under pharmacological or exertional challenge. If cardiac and respiratory sensitivity correlate under isoproterenol or exercise and not at rest, domain generality survives as a state-dependent property, which is a more interesting claim than the one being defended.
- Three or more channels in one sample with matched psychophysics. Ferentzi covered breadth with weak instruments; Banellis covered instruments with two channels. Gastric is the notable absence, since it is where coherence was originally claimed — and Levakov et al. (2023) complicates the obvious way of filling it. The gastric channel’s best-developed measure is a coupling quantity (gastric-network phase synchrony), which has no test–retest reliability in individuals; the reliable gastric measurement is the EGG signal itself, which is physiology and not perception. So gastric cannot currently enter this debate as a third psychophysical axis at all. Someone would first have to build a gastric detection task with modern psychophysics, and the last serious attempts date to the 1980s. Updated with Haruki & Ogawa (2023): the gastric channel now has a neural signature obtained without any instrumentation — cue the word STOMACH and a distinctive cortical profile appears. That is enough to enter the substrate argument above and not enough for the psychophysical one, since the paradigm has no correct answer. The gap has narrowed in the wrong currency.
- Decomposing the confidence correlation. Does cross-modal confidence still bind after scale-use variance is modelled out? A generative model of interoceptive metacognition is the authors’ own proposal.
- A control-side test. Whether interoceptive-control is channel-specific is entirely unknown, and cannot be inferred from these perceptual nulls. There is currently only one channel with a control task (respiratory-tracking-task), so this question may be unaskable for some time.
Why open rather than resolved
Because the null is moderate rather than decisive, because it is confined to two channels at rest in one demographic, and because the most plausible defence of domain generality — that it appears only under perturbation — is not merely unrefuted but untested. What is resolved is narrower and still consequential: there is no demonstrated person-level interoceptive ability, and claims of the form “X has poor interoception” are, on present evidence, claims about an organ.