Interoceptive taxonomy (Farb et al. Box 1)
Proposed in Farb et al. (2015), Box 1, following Garfinkel & Critchley (2013). Motivation: treating interoception as a single construct produces spurious inferences — e.g., meditators’ strong interoceptive attention tendencies do not predict superior heartbeat-detection accuracy (Khalsa et al. 2008; Parkin et al. 2013), a finding that only makes sense once attention and accuracy are separated. See does-mindfulness-enhance-interoceptive-accuracy.
The seven constructs
- Awareness — the most common measure (e.g., sense of one’s heart beating), usually operationalized as reportability; limited because interoceptive processes can operate implicitly (e.g., thermoregulatory shivering during sleep).
- Coherence — degree to which objectively observable interoceptive signals manifest in reportable experience (e.g., hypoglycemia can promote irritability without awareness of low blood sugar). Varies widely between individuals.
- Attention tendency — whether a person habitually attends to particular interoceptive signals (or to interoceptive vs. exteroceptive information generally).
- Sensitivity — the minimum threshold for detecting interoceptive signal change; may operate at multiple representational levels culminating in conscious access. Its signal-detection-theory counterpart, specificity, is rarely used in interoception research.
- Accuracy — ability to reliably discriminate interoceptive signals from noise/competing signals and between intensity levels; the most commonly investigated measure (usually via the heartbeat-detection-task); a function of sensitivity and specificity; improvable with training.
- Sensibility — an individual’s personal account of their internal-sensation experience, including confidence in their own interoceptive ability and feelings of engagement; gauged via interview/questionnaire (e.g., Porges Body Perception Questionnaire — though its interoceptive validity is questioned since it may largely index anxiety; also the Scale of Body Connection and the maia).
- Regulation — how well a person matches an interoceptive signal to a desired state, whether via reappraisal/suppression/distraction (shaping the signal) or contemplative acceptance/curious examination (shaping the relationship to the signal, without necessarily changing it).
The taxonomy in use
Oldroyd et al. (2019) give the sensibility/accuracy distinction a concrete payoff. Attachment anxiety correlated with higher maia Noticing and Emotional awareness but sharply lower Not-worrying (r = -0.43) — a person who attends to the body constantly and is frightened by what they find. Collapsed to a single “interoceptive awareness” score this profile reads as good interoception; only the multidimensional measure reveals it as hypervigilance. Note also what the taxonomy withholds: because the MAIA indexes sensibility alone, none of their correlations bear on interoceptive accuracy, and the paper’s title-level claim about “interoception” is really a claim about sensibility. Their Study 2 congruence measure (self-report-physiology-congruence) is a rare attempt to operationalize coherence directly.
Terminological tension with Garfinkel et al. (2015)
Garfinkel et al. (2015) combine sensitivity and accuracy into a single interoceptive accuracy (objective task performance), and define interoceptive awareness as “metacognitive awareness of interoceptive accuracy, e.g., confidence–accuracy correspondence.” Farb et al. explicitly push back on this: “metacognitive awareness” already has an established, different meaning in contemplative practice — the capacity to take awareness/thought itself as an object of attention (per Smallwood & Schooler 2006; Herbert & Forman 2011). To avoid conflating these, Farb et al. propose the term coherence for confidence–accuracy correspondence instead, reserving “metacognitive awareness” for its contemplative sense. This is a genuine, acknowledged terminology conflict between near-contemporaneous papers, not a simple restatement — see the note on interoceptive-sensitivity.
A third taxonomy: Khalsa et al. (2018), Table 2
The field’s consensus roadmap offers its own eight-feature nomenclature — for interoceptive awareness, which it, like Farb et al., regards as an umbrella term stretched too far (“first used to describe a self-report subscale… it has subsequently been used to encompass any (or all) of the different interoception features”). The features:
- Attention — observing internal body sensations.
- Detection — presence or absence of a conscious report.
- Magnitude — perceived intensity.
- Discrimination — localizing a sensation to a channel/organ and telling it from others.
- Accuracy (Sensitivity) — correct and precise monitoring. (Note: Khalsa equates accuracy and sensitivity here, where Farb et al. hold them apart.)
- Insight — metacognitive evaluation of one’s own performance (e.g. confidence–accuracy correspondence).
- Sensibility — self-perceived tendency to focus on interoceptive stimuli (a trait measure).
- Self-report (Scales) — psychometric assessment via questionnaire.
The roadmap frames the same core commitment as the other two: distinguish sensation (raw signal) from perception (what is made of it), and stop treating “interoceptive awareness” as one thing.
Garfinkel’s three dimensions, stated first-hand (Quadt et al. 2018)
The wiki reconstructs the Garfinkel et al. (2015) framework above mostly from Farb et al.’s pushback against it. Quadt, Critchley & Garfinkel (2018) state it directly, from Garfinkel’s own group, and the wording is worth pinning down because the wiki cites these terms constantly:
- Interoceptive accuracy — objective performance on a behavioural test (how accurately one performs a heartbeat-tracking task).
- Interoceptive sensibility — subjective belief about one’s own ability to perceive bodily signals (questionnaire, e.g. the Body Perception Questionnaire, or rated confidence in one’s task performance).
- Metacognitive interoceptive awareness — the insight linking the two, derived from confidence–accuracy correspondence; the review argues this is “a most appropriate use of the word ‘awareness’” — the direct rebuttal of Farb et al.’s objection to that usage.
Quadt et al. also embed these three “psychological” dimensions in a larger scheme the other taxonomies don’t foreground: a first dimension for the afferent signal itself (measured by neuroimaging change or the heartbeat-evoked potential, whose amplitude correlates with detection ability — the wiki’s cleanest neural index of the raw interoceptive signal, distinct from any perceptual report); a second dimension for the signal’s influence on other cognition/behaviour without requiring awareness (cardiac-timing effects on decision, emotion, memory); and a further executive dimension — the capacity to flexibly attend to, use, and switch between interoceptive and exteroceptive representations. So the Sussex scheme runs signal → influence → accuracy/sensibility/awareness → executive control, a fuller stack than either Farb’s seven flat constructs or Khalsa’s eight.
A fourth list, and the reason taxonomies keep multiplying (Berntson & Khalsa 2021)
Berntson & Khalsa (2021) — Khalsa again, three years on — give awareness seven features in their Figure 1: detection, attention, insight, magnitude, discrimination, accuracy, sensibility. Against his own 2018 eight, “self-report (scales)” has dropped out. Nothing turns on the difference, and that is the point: the same author produces a slightly different list each time because the list is a description of a research literature, not a partition of a mechanism.
Two things this review adds that the three taxonomies above do not.
Awareness is the last stage of five, and the rarest. The paper’s continuum runs sensors → pathways → networks → circuits → awareness, and “most processing of interoceptive signals occurs beyond the conscious awareness of the organism.” Conscious interoception arises mainly on perturbation (quickened breath, surging heartbeat, full bladder) or threat (angina, nausea, kidney stone); some organs — liver, kidney — are never felt at all. So every construct on this page is a way of carving up the thin conscious tip of a system whose business is conducted below it. There are also prominent clinical individual differences in what reaches awareness at all: the review’s example is sex differences in the detection of angina during myocardial infarction, which is a taxonomy problem with mortality attached.
Sensitivity does not generalize across channels. The sharpest methodological item: “perceptual sensitivity for one signal may not generalize to others” (Ferentzi et al. 2018 — Multichannel investigation of interoception: sensitivity is not a generalizable feature). All three taxonomies above are written as though “accuracy” and “sensibility” name person-level traits. If they are channel-specific, then a cardiac accuracy score is not an interoceptive accuracy score, and much of the wiki’s cross-study comparison is comparing different quantities under one name. This is the same worry is-the-heartbeat-counting-task-valid pursues from the validity side, arriving here from the generalizability side.
This is no longer a caution held at one remove. Banellis et al. (2026) tested it directly in the axis pair Ferentzi et al. omitted, with matched psychophysics: cardiac and respiratory sensitivity, precision and metacognitive efficiency are uncorrelated at N = 241, with moderate Bayesian evidence for the null. The taxonomies’ shared assumption — that these words name properties of a person — is the thing that failed. See is-interoception-domain-general, created with that ingest, and the section below on what survived the test.
Related: interoceptive processing can influence exteroceptive perception (Motyka et al. 2019), so the constructs are not cleanly separable from exteroceptive ones at measurement either.
And there is no animal version of any of this. “There are currently no animal models of interoceptive awareness,” though insular homology between humans and non-human primates may make one worth attempting. Every construct on this page is defined by a report. See can-we-know-animal-feelings.
A fifth list, stated first-hand by its owner (Greenwood & Garfinkel 2025)
The Desmedt section below distinguishes Desmedt et al.’s hierarchy-by-specificity from “the otherwise-similar Suksasilp & Garfinkel (2022) hierarchy organised by processing depth.” That Suksasilp & Garfinkel scheme is the one the wiki had only named. Greenwood & Garfinkel (2025) lay it out in full — an eight-dimension model running from the raw signal to its interpretation:
- Nature of afferent signals · 2. Neural representation of afferent signals · 3. Preconscious impact of afferent signals · 4. Interoceptive accuracy · 5. Self-report and interoceptive beliefs · 6. Interoceptive insight · 7. Attention to interoceptive sensations · 8. Attribution of interoceptive sensations.
Two things place it against the four lists above.
It is Garfinkel’s three, re-embedded — so the “single measure is three” line becomes eight. Dimensions 4/5/6 are her accuracy / sensibility / awareness verbatim (Quadt et al. restated). What the eight-list adds is to promote the signal (dim. 1–2) and its preconscious influence (dim. 3, the cardiac-timing level) to dimensions in their own right, and to split the higher-order end into attention and attribution — the latter absorbing Schachter–Singer misattribution and alexithymia as a taxonomy entry, which none of the other four do explicitly. So the Sussex scheme is the only one of the five to carry an attribution dimension, and the only one built to run each level straight into an emotion example (Greenwood & Garfinkel’s Table 2).
It inherits, and states, the assumption Desmedt below predicts will bite. Read as person-level traits, dimensions 4–8 are the exact quantities Banellis et al. found do not travel between organs, and dimension 4 is the cardiac-accuracy measure the review’s clinical sections rest on. Greenwood & Garfinkel’s own Limitations concede the cardiac-domain narrowness — so the fifth taxonomy is the first whose authors flag, in the same paper, the organ-specificity crack the other four leave implicit. It changes nothing about the collisions catalogued on this page; it adds a fifth partition of the same conscious tip, from the group whose three-way split the whole page is organized around.
The gap the taxonomies cannot paper over: non-verbal populations (Chen et al. 2021)
Every construct on this page above is operationalized through a verbal report of a counted or rated quantity — heartbeats counted, confidence rated, a questionnaire item endorsed. Chen et al. (2021) state the consequence as the field’s lifespan bottleneck: there is no measure of interoceptive sensitivity suitable for non-verbal populations such as infants, or older adults with dementia.
That single sentence explains why two of this wiki’s threads are thin at exactly the ends where the interesting claims live:
- The developmental end. social-origins-of-interoception and Oldroyd et al. make claims about how interoception is built in infancy. The evidence base is one cited attempt at neurobehavioural infant measurement (Maister et al. 2017). Tsakiris, writing the developmental section of Quigley et al., says the emergence of interoceptive awareness in early life “remains unexplored” and blames limited-range measures in children. Same diagnosis, two papers, one issue.
- The neurodegenerative end. age-related-interoceptive-decline runs out of measurement precisely where Bonaz et al. report Alzheimer’s impairment across all four taxonomy dimensions — a finding that requires exactly the instrument Chen et al. say does not exist for that population.
The NIH review pairs this with a portfolio-level version of the same point: “objective and quantitative assessments of interoception” is one of its Outstanding Questions, and its inventory of what human research can currently do — heartbeat measures, skin conductance, self-report scales — is summarized as “largely limited to a handful of approaches and remain mostly correlational.” microneurography is nominated as the one recent addition offering objectivity, for the second time in this journal issue.
So the taxonomies are not merely multiplying; they are all partitions of a construct measurable in one kind of participant. See interoception-exteroception-boundary for the parallel problem on the other axis — the constructs are also channel-specific, per Ferentzi et al. above.
The one construct nobody measures: regulation (Allen 2026)
Read the four lists on this page for what they have in common rather than where they collide, and a structural fact appears that none of the source papers remark on: almost every construct across all four taxonomies is afferent.
Detection, attention, magnitude, discrimination, accuracy, sensitivity, sensibility, insight, coherence, awareness — all of them name something the brain reads. Exactly one entry, in one of the four lists, names something the person does: Farb et al.’s regulation, defined above as “how well a person matches an interoceptive signal to a desired state.” Garfinkel’s three have no efferent term. Khalsa’s eight have none. Berntson’s seven have none.
Allen (2026) opens by naming the consequence: “Interoception research has focused primarily on perception, with comparatively little attention to the control of bodily state.” The mismatch is with this wiki’s theory pages, which are about control throughout — active-inference resolves prediction error by changing the body, allostasis is anticipatory regulation, Chen et al. fold descending regulation into the definition of interoception complete with “interoceptive effectors,” and Berntson & Khalsa insist the unit is a circuit because afference cannot be described apart from the efferent systems that change what is sensed. The models are closed-loop; the measurements are afferent-only.
The reason is not oversight, and it is the same reason the taxonomies are channel-specific: most interoceptive channels cannot be voluntarily driven at all, so there is nothing to score. Regulation stayed a definition without a task for a decade because a task requires a channel the participant can act on — and there is essentially one, respiration. See interoceptive-control for what that restriction costs the construct, and respiratory-tracking-task for the first instrument to occupy the slot.
Which sharpens the accuracy-vs-sensitivity quarrel below into something less parochial: the field has produced four competing partitions of the perceptual half of interoception while leaving the other half with one word and no instrument.
The other missing half: use, not just regulation (Murphy 2022)
The section above finds one non-perceptual construct — regulation — and notes the taxonomies are otherwise all afferent. Murphy (2022), commenting on the fifth list (Suksasilp & Garfinkel 2022), names a second non-perceptual construct the lists also omit, and it is not efferent: propensity to use interoceptive signals — whether a person actually deploys a signal they can perceive, versus relying on external cues.
This is a different axis from every entry on this page. Detection, attention, magnitude, accuracy, sensibility, insight all ask how well is the signal read. Regulation asks how well is the body driven to a state. Propensity asks is the signal consulted at all — internal versus external cue weighting, which the interoception-exteroception-boundary page treats as a definitional problem and Murphy recasts as a person-level disposition, further split by which process (a person may weight the body heavily for hunger, lightly for emotion) and by whether the use is adaptive.
Two reasons it belongs on this page rather than beside it.
It supplies a name for the disposition the page’s anomalies keep circling. This page records, repeatedly, that a belief/attention component dissociates from ability and often generalizes out of the body — confidence travelling to auditory tones, sensibility splitting into believed-accuracy and deployed-attention, meditators attending without detecting. Those are all facts about use wearing the vocabulary of perception. Murphy’s construct is the one that names the thing: the taxonomies bundle “do you consult the body” into “can you read the body,” and the two come apart.
Its motivating datum is a sex difference the perceptual taxonomies cannot absorb. Men outperform women on lab interoceptive accuracy, the gap vanishes in naturalistic settings with external cues available, and emotion relates to interoception more strongly in men — yet women are better at emotion. If ability drove use, the better-emotion group would read the body better. Murphy’s reading: the groups differ in propensity, and for some processes weighting external cues is the superior policy. See is-more-interoceptive-awareness-better and interoceptive-propensity.
Like regulation, propensity is theorised and unmeasured — there is no instrument, and its trait-vs-state status is open (state-vs-trait-interoception). The page now records two non-perceptual axes the five taxonomies leave out, for the same structural reason: the field’s instruments read signals, and neither using a signal nor acting on a channel is a signal to read.
The one construct that did travel, and it is not interoceptive (Banellis et al. 2026)
Every list on this page assumes its constructs are properties of a person. Banellis et al. (2026) put four of them — sensitivity, precision, metacognitive bias, metacognitive efficiency — through the same estimation pipeline in the cardiac, respiratory and auditory modalities at once. The result cuts the taxonomies in a place none of them anticipated.
Nothing performance-related travelled. Cardiac and respiratory sensitivity, precision and metacognitive efficiency were all uncorrelated (BF01 = 6.39, 6.30, 5.04). Ferentzi’s warning, confirmed on better instruments.
One thing travelled strongly: metacognitive bias — mean confidence. Cardiac↔respiratory r = 0.510. And the largest correlation in the whole study was respiratory↔auditory confidence, r = 0.642 — an interoceptive judgment binding more tightly to a judgment about tones than to the other interoceptive judgment.
That single comparison is the most consequential thing this ingest does to this page. The three taxonomies each contain a confidence-flavoured construct — Garfinkel’s awareness, Farb’s coherence, Khalsa’s insight — and each treats it as interoceptive. Split into its two components, the interoceptive-specific component (efficiency) is channel-bound, and the component that generalizes (bias) generalizes right out of the body. So the taxonomies were not merely giving one thing three names; they were bundling a domain-general disposition with an organ-specific competence and naming the bundle after the body.
The practical consequence for reading this wiki: sensibility measures — the maia, the BPQ, rated confidence — sit on the domain-general side of that split. Which supplies one mechanism for a pattern recorded independently on four pages: sensibility and accuracy keep failing to correlate (interoceptive-sensitivity), meditators attend more without detecting better (does-mindfulness-enhance-interoceptive-accuracy), and attachment-anxious profiles notice more while worrying more. Four anomalies, one dissociation.
The construct with an instrument nobody scores: attention (Haruki & Ogawa 2023)
Every list on this page contains attention — Farb et al.’s “attention tendency,” Khalsa et al.’s “attention,” Berntson & Khalsa’s “attention” — and the wiki has generally met it as a trait inferred from questionnaires (maia Noticing) or from what meditators say they do.
It has an instrument, and the wiki had never named it. The interoceptive-attention-task — cue an organ, hold attention on it for ten seconds, subtract an exteroceptive control — is the dominant paradigm in interoceptive neuroimaging, standing behind results cited across this wiki on farb-2015-interoception-contemplative-health, mindfulness-meditation, weng-2021-interventions-of-interoception and quadt-2018-interoception-health-disease. Created as a page with the Haruki & Ogawa (2023) ingest, the first study in raw/ to run it on two organs at once.
Two things it changes here.
Attention is the one construct measurable without a report of a quantity — and it still needs a report. The non-verbal measurement gap recorded above turns on every construct being operationalized through a counted or rated number. Attention looks like the exception: the dependent variable is brain activation, and the participant supplies nothing. But compliance is unverifiable — nothing distinguishes attending to your stomach from being told to and thinking about something else — so the paradigm is verified by exactly the self-report it appeared to escape. The route out is Weng’s decoding approach, where a classifier reads interoceptive attention off the brain and can in principle check the instruction rather than trusting it.
And the constructs are channel-specific at the neural level, not only the behavioural one. The section above records Banellis et al. showing that sensitivity, precision and efficiency do not travel between organs. Haruki & Ogawa show that attention does not either, in a different sense: attending to the heart and attending to the stomach are different brain events in the same person, differing in the anterior insula and dissociable by pattern classifier in the middle insula. Two constructs from these lists, two evidence classes, same verdict — the taxonomies name channel-bound capacities and the words hide it. See is-interoception-domain-general.
A terminological consequence worth flagging, because this paradigm’s papers commit it consistently: Haruki & Ogawa’s conditions are labelled “cardiac interoceptive awareness” and “gastric interoceptive awareness” on the strength of attention having been deployed. On every taxonomy on this page, attention and awareness are separate entries — and the wiki’s clearest evidence that they dissociate is that meditators have the first without gaining the second (does-mindfulness-enhance-interoceptive-accuracy). A paradigm that measures one construct and is titled after another is the exact failure these taxonomies were written to prevent.
The levels dissociate within one channel too, and they rank (Harrison et al. 2021)
Everything above tests the taxonomies across channels — do these words name person-level properties that travel from heart to lung? Harrison et al. (2021) stay inside one channel and ask a different question: do the levels dissociate from each other, and does an affective trait attach to all of them equally?
Sixteen measures in 60 people, all respiratory or affective: four affective questionnaires, four interoceptive questionnaires (MAIA, BPQ, breathing catastrophizing, breathing vigilance), four FDT parameters (threshold, decision bias, metacognitive bias, M-Ratio), and four brain measures.
They dissociate. Cross-task correlations in the 16-measure matrix were sparse. Within the questionnaire block, correlations were strong and pervasive; between blocks, weak and scattered. The authors read this as supporting “potentially separable levels of breathing-related interoception, as proposed” — citing Critchley & Garfinkel (2017). So the taxonomies’ central commitment survives a within-channel test, which is worth stating because the cross-channel tests above are usually read as bad news for them.
But they rank, and the ranking is uncomfortable. The paper’s PCA orders the measures by how much of the anxiety-group difference each carries:
- affective questionnaires (depression, state anxiety, ASI, GAD-7)
- breathing catastrophizing and negative MAIA awareness — i.e. sensibility
- metacognitive bias, body perception, metacognitive efficiency — insight/coherence/awareness
- perceptual threshold, decision bias — sensitivity/accuracy
- anterior insula activity — last
Read against the four lists on this page, that is an ordering by distance from the body. The construct every taxonomy treats as the field’s gold standard — objective task performance — carries second-least. The constructs the taxonomies are most sceptical of, because they are “only” self-report, carry most.
Two consequences worth recording:
It gives the sensibility/accuracy split a direction, not just a distinction. The taxonomies established that sensibility ≠ accuracy. The Banellis section above added that the two sit on opposite sides of a domain-general/organ-specific divide. This adds that in the one place both have been measured against a clinically relevant trait in one sample, sensibility wins. The field’s measurement effort and its explanatory payoff point in opposite directions.
The deflationary reading is available and does not help much. Higher levels may simply be less noisily measured, so the ordering could reflect detectability. The authors say so themselves. But that concession leaves the practical conclusion untouched and makes it worse for the accuracy programme: if a psychophysical threshold is too noisy to detect a relationship this large, the instrument problem is the finding.
The assumption none of the four lists states: that these are properties of a person at a time (Wallman-Jones et al. 2023)
The sections above dismantle the taxonomies from two directions — the constructs do not travel between organs (Banellis, Haruki & Ogawa), and they do not rank the way the lists imply (Harrison). Wallman-Jones et al. (2023) add a third axis the lists are equally silent on: time.
Measuring self-reported interoception ten times a day for a week in 70 people, they report an ICC of 0.51 — barely half the variance is between people. The rest moves within them, on a fifteen-minute timescale, with whether they had been exercising or looking at a screen. Every construct on this page is written as a noun a person has. In the one case where anyone has partitioned the variance, half of it belongs to the afternoon. See state-vs-trait-interoception.
And a datum that runs against this page’s most reliable pattern. The sections above record accuracy and sensibility failing to correlate so consistently that the failure has become the wiki’s explanation for four separate anomalies. Here baseline heartbeat-counting accuracy positively predicted week-averaged self-reported interoception (B = 4.50, β = 0.21, p = .007). Recorded as friction rather than refutation — the sensibility measure is a state scale aggregated over a week rather than a trait questionnaire, N = 70, and the confidence interval is wide (1.31–7.69) — but it is the wiki’s first positive accuracy↔sensibility association and it should not be quietly filed with the nulls. If it holds, it suggests the dissociation may be partly an artefact of comparing a single-session accuracy score against a trait questionnaire, i.e. a state/trait mismatch rather than a construct dissociation.
And the entry every list shares is itself two constructs (Ventura-Bort et al. 2021)
The sections above attack the taxonomies from outside: the constructs do not travel between organs, do not rank as the lists imply, and are not stable across an afternoon. Ventura-Bort, Wendt & Weymar (2021) attack from inside, on the one entry that appears — under that exact name — in all four lists: sensibility.
Eight self-report instruments in 109 people, four about the body (ICQ, IAS, MAIA-2, and BPQ-adjacent material) and four about emotions (TAS-20, MAS, RDEES, TMMS). A PCA forced to two components split them, and not along the body/emotion line the design assumed:
- Factor 1, “Sensibility” (38.9% of variance) — believed accuracy: IAS, MAIA Attention regulation and Trusting, reversed ICQ, reversed TAS-20 (all three subscales), MAS Labeling, RDEES Differentiation, TMMS Clarity and Repair.
- Factor 2, “Monitoring” (10%) — deployed attention: MAIA Noticing and Emotional awareness, MAS Monitoring, RDEES Range, TMMS Attention.
Now read the four lists’ definitions of sensibility against that:
| taxonomy | its “sensibility” | which factor |
|---|---|---|
| Garfinkel / Quadt et al. | ”subjective belief about one’s own ability to perceive bodily signals” | Sensibility |
| Khalsa et al. (2018) | “self-perceived tendency to focus on interoceptive stimuli” | Monitoring |
| Berntson & Khalsa (2021) | same | Monitoring |
| Farb et al. (Box 1) | “personal account of internal-sensation experience, including confidence in their own interoceptive ability and feelings of engagement” | both, conjoined |
So the three fixes for the field’s terminological chaos are, on this one entry, naming two different things with one word — and Farb’s version bundles them in a single definition. This is not the taxonomies disagreeing about a boundary; it is the same word picking out orthogonal factors in different papers.
Why it matters more than a labelling complaint: the two halves have opposite outcomes. Sensibility correlated r = .58 with well-being; Monitoring r = .12 (Z = 3.94). On MAIA Not-worrying the signs invert: r = .40 vs −.20. And Sensibility predicted higher granularity for negative emotions where Monitoring predicted lower granularity and lower felt intensity. A study that cites “sensibility” as a predictor is therefore not merely imprecise; depending on which instrument it used, it may have measured the protective half or the costly one.
Three things this settles, or nearly.
It supplies a mechanism for the page’s recurring anomaly. This page records four separate cases of sensibility behaving oddly against accuracy — meditators attending without detecting, attachment-anxious noticing-with-worrying, confidence generalizing out of the body, sensibility outranking accuracy against anxiety. If “sensibility” is a mixture of a confidence construct and an attention construct in proportions that vary by instrument, then studies using different questionnaires were measuring different mixtures, and some of the inconsistency is compositional rather than substantive.
It refines the Harrison ranking rather than contradicting it. Harrison et al. rank sensibility above accuracy for carrying affective variance. This says only one half of sensibility carries the well-being variance — and it is the belief half, not the attention half. Note the convergence in detail: Harrison’s second-ranked measures were breathing catastrophizing and negative MAIA awareness, which are the worry-flavoured items, and here the worry-flavoured pole is the one that loads with Monitoring and predicts poorly.
It does not rescue the taxonomies by being tidier than they are. Two brakes. Only five MAIA subscales were entered as predictors — the other three (Not-distracting, Not-worrying, Self-regulation) were used as outcomes, so the instrument was partitioned by the analysts before the factors were extracted. And the two-factor solution was forced: four components had eigenvalues > 1, and the authors imposed two because the four-component solution loaded by questionnaire. The finding is real and the shape is interpretable, but the number two is a choice.
And the largest scope claim. The factors sort bodily and emotional self-report instruments together. Whatever these questionnaires measure, it is not specific to interoception — it is a disposition toward one’s own internal states, of which the bodily items are a sample. That is the Banellis verdict (the generalizing component is not interoceptive) reached from the questionnaire side rather than the psychophysics side, and it is why alexithymia’s TAS-20 and the IAS end up on one factor. See is-interoception-domain-general.
Three taxonomies, and where they collide
The field now has three overlapping nomenclatures — Farb et al.’s seven, Garfinkel et al.’s three, and Khalsa et al.’s eight — and they do not map cleanly onto one another. The sharpest non-alignment is over accuracy vs sensitivity:
- Farb et al. keep sensitivity (detection threshold) and accuracy (discrimination from noise) as separate constructs.
- Khalsa et al. fuse them into one feature, “Accuracy (Sensitivity).”
And over the metacognitive term:
- Garfinkel calls confidence–accuracy correspondence interoceptive awareness; Farb calls it coherence (reserving “awareness” for reportability and “metacognitive awareness” for the contemplative sense); Khalsa calls it insight.
So the same confidence–accuracy quantity carries three different names across three taxonomies published within three years by overlapping author sets (Critchley is on both the Garfinkel and Khalsa lists; see hugo-critchley). This is not a failure of any one paper but the exact terminological chaos each of them was written to fix — and the fact that three fixes coexist is itself the state of the field. A reader porting a claim from one paper to another must check which taxonomy’s “accuracy,” “awareness,” or “sensibility” is meant.
The meta-diagnosis: broad constructs guarantee low convergence (Desmedt et al. 2023)
Every section above records a specific non-convergence — sensibility ≠ accuracy, cardiac ≠ respiratory, attention ≠ awareness, sensibility is two things. Desmedt, Luminet, Maurage & Corneille (2023) supply the reason the wiki had reached empirically without stating: the broader a construct, the more heterogeneous the phenomena it covers, so the lower the convergence between measures assigned to it must be. The taxonomies were built to standardise the field’s terminology by proposing broad dimensions — and breadth is precisely the property that produces the divergence documented across this page. The three-taxonomies problem is not an accident of drafting; it is what happens when you name a heterogeneous measurement literature with a small number of superordinate terms.
Two numbers this review adds to the page’s convergence tally:
- Within the cardiac channel, not just across channels. Hickman et al.’s (2020) meta-analysis of 22 studies found the HCT and the Heartbeat Discrimination Task — both called “interoceptive accuracy,” same organ — share 4.4% of variance (r = 0.21). The Ferentzi cross-channel null above has a within-channel twin.
- The statistical consequence of naming them alike. Carlson & Herdman (2012): if two measures correlate r = 0.30 and one correlates r = 0.30 with an outcome, the other’s correlation with that outcome can lie anywhere from −0.95 to 0.95. So labelling weakly-correlated measures with one construct name makes cross-measure inconsistency the expected result — a manufactured “replication crisis.” This is the formal engine behind is-interoception-domain-general’s warning that cardiac conclusions do not license interoceptive ones.
A sixth partition, and the only one shaped like a design rather than a list (Murphy et al. 2020). Everything above this line is a flat enumeration of constructs — seven, three, eight, seven, eight — and Desmedt’s contribution is to replace flatness with hierarchy. Murphy, Brewer, Plans, Khalsa, Catmur & Bird (2020) replace it with a 2×2 factorial: what is measured (accuracy of perception vs the body being the object of attention) crossed with how (objective performance vs self-reported belief). Four cells, and — unlike any list on this page — a prediction that can fail: measures should relate when they share the what, whichever the how. They built the missing cell’s instrument (the IAS, self-reported accuracy) precisely so the prediction could be tested, and the test is the section below.
The constructive alternative: a hierarchy by specificity, not a flat list. Desmedt et al. propose replacing the competing flat taxonomies with a hierarchical framework organised by level of specificity. Interoception at the top; below it broad factors — interoceptive attention, sensing, interpretation, memory (the first scheme on this page to carry a memory factor and to include non-conscious processing as a matter of definition); below those, subfactors onto which the old dimensions map (Garfinkel’s accuracy → “interoceptive detection”; Khalsa’s discrimination → “interoceptive localization”; plus interpretation-side constructs the flat lists omit — somatosensory amplification, interoceptive worrying, emotional awareness); and at the bottom, measure-tied subfactors where a construct has exactly one measure and claims no more (the Schandry HCT becomes “the capacity to estimate heart rate via mental counting,” full stop). The organising principle is specificity, which distinguishes it from the otherwise-similar Suksasilp & Garfinkel (2022) hierarchy organised by processing depth. The scheme is offered as illustrative — the number of levels and the content of the factors both await construct-validity work (Clark & Watson 2019) — but its rule for reading this wiki is immediate: a claim is only as broad as the measure that grounds it. That is the discipline the “accuracy,” “sensibility” and “awareness” collisions above have been quietly demanding.
The distinction earns its keep at last, and it puts a boundary on this page’s favourite null (Murphy et al. 2020)
The Ventura-Bort section above splits sensibility into believed accuracy and deployed attention, and closes on what it could not do: “Untested: Ventura-Bort et al. administered no accuracy task at all.” Murphy et al. (2020) ran the missing half, a year earlier, with the IAS that Ventura-Bort et al. later used.
Self-report splits along the what axis, replicated three times. The two believed-accuracy scales (IAS, reverse-scored ICQ) correlate at r = .57–.71; the attention scale (BPQ) correlates with neither (|r| ≤ .09 in five of six estimates), and every accuracy↔accuracy correlation was significantly larger than its accuracy↔attention counterpart. Ventura-Bort’s forced two-component PCA and this correlation structure are two methods, two labs, two countries, one seam.
And the how axis crosses successfully. Heartbeat-counting accuracy correlated with the IAS (r = .271, n = 52; r = .336, n = 33) and with nothing else in the battery — not the BPQ, not the ICQ, not the TAS-20, not time-estimation performance. State confidence ratings behaved the same way: correlated with the IAS and with task performance, uncorrelated with the BPQ.
What this does to the rest of the page is larger than the effect size. Six sections above record self-report failing to track objective accuracy, and the wiki has read that as evidence against the taxonomies — words naming things that come apart. Look at what was in the self-report slot each time: maia Noticing, the BPQ, generic body-awareness scales. Attention measures, every one. The null is real and it is a null about attention. Where the what is matched, the correlation appears.
That is the first time on this page a taxonomic distinction has predicted rather than merely described — the lists have been diagnoses of a messy literature, and this is one of them making a wrong-able claim and surviving. It also reframes rather than removes the trouble: the taxonomies’ problem is not that their constructs fail to behave, it is that the field’s instruments are unevenly distributed across the cells, so most published comparisons are mismatched by construction.
Two brakes, and they are not small. r = .27 and r = .34 at n = 52 and n = 33, both at p ≈ .047 — about 7–11% shared variance, in samples too small to separate much. And the objective anchor is the heartbeat counting task: if that score is largely guessing from a believed heart rate, then its correlation with a questionnaire asking “can you accurately perceive your heartbeat?” is a belief↔belief correlation wearing a how axis it has not earned. Held as the strongest available evidence for the distinction, at the weight two small correlations can bear.