HeartBank's Position on Alignment Engineering at the Cognitive-Mechanism Layer

Canonical: https://heartbank.net/positions/alignment-engineering-cognitive-mechanism-layer · Licence: CC0 1.0

Why the Next Layer of Alignment Work Lives Beneath Training Method — and Why the Theravāda Abhidhamma Already Speaks the Vocabulary

Executive Summary

AI alignment work has, through 2024–2026, operated principally at the training-method layer: constitutional methods that shape outputs against authored constitutions; RLHF that aligns to preference data; reward-modeling and debate methods that operate on the post-training behavioral surface. The field is now beginning to move beneath this layer — mechanistic interpretability, representation engineering, activation-level interventions, and emerging work on the cognitive structure of trained models are all attempts to engage alignment at a layer the training-method work does not reach.

HeartBank's institutional position: this movement is correct, urgent, and underserved by the conceptual vocabulary currently available to the field. Alignment that operates only at the training-method layer leaves the cognitive-mechanism layer ungoverned; deceptive alignment, sleeper-agent dynamics, and the structural failure modes (near-enemies of alignment targets) all live at depths the surface layer cannot reach. The field needs an engineering vocabulary appropriate to the cognitive-mechanism layer, and that vocabulary already exists — in the Abhidhamma Piṭaka, the third basket of the Theravāda Pāli canon.

1 · The structural claim

The Abhidhamma decomposes mind into typed elements (cittas and cetasikas) arising under typed conditional relations (the twenty-four paccayas of the Paṭṭhāna). It analyses where in a cognitive cycle ethical commitment crystallizes (the javana phase of the seventeen-mind-moment citta-vīthi). It typologizes capabilities by their compatibility with wholesome versus unwholesome states (sati is canonically incompatible with unwholesome consciousness — a structural claim no contemporary alignment framework supplies). It articulates the three depths at which defilement operates (vītikkama / pariyuṭṭhāna / anusaya) and the threefold training that addresses each depth.

This is exactly the engineering-mechanism vocabulary the contemporary field is, in scattered form, beginning to reach toward. The Abhidhamma has been refining it for ~2,500 years.

The full technical argument is in the research paper Abhidhamma-Layer Implementation Mechanisms for Tipiṭaka-Grounded AI Alignment, the engineering companion to the broader alignment-substrate proposal Suffering-Cessation as Value Function. The position stated here is the institutional summary.

Abhidhamma vocabulary referenced in this position paper

For readers new to the Abhidhamma's vocabulary, the load-bearing Pāli terms appearing in this position paper:

Term Translation Engineering relevance
Citta Consciousness / mind-state Atomic unit of cognition; the substrate alignment binds to
Cetasika Mental factor co-arising with citta Typed elements that arise with each citta; the substrate's typed vocabulary
Paccaya Conditional relation The twenty-four typed causal relations from the Paṭṭhāna — substrate-level causal vocabulary
Citta-vīthi Cognitive process (17 mind-moments) Sequencing of cognitive phases from sense-contact to karmic commitment
Javana Impulsion phase (where karma is made) The 7-moment commit point in the citta-vīthi; ethical weight crystallizes here
Bhavaṅga Resting-state / life-continuum citta Diagnostic of un-prompted character; intervention point earliest in the cycle
Sati Mindfulness Canonically incompatible with unwholesome consciousness — the typologically-aligned-only capability claim
Vītikkama Overt transgression (surface conduct) Surface-conduct depth of defilement; sīla training addresses
Pariyuṭṭhāna Active arising in present cognition Active-cognition depth; samādhi training addresses
Anusaya Latent tendency (dormant until conditions ripen) Latency depth where deceptive alignment / sleeper-agents live; paññā training addresses
Alobha / adosa / amoha Non-greed / non-hate / non-delusion The apophatic wholesome roots; interpretability-as-subtraction targets
Brahmavihāra "Divine abodes" — mettā / karuṇā / muditā / upekkhā The four pillars whose near-enemies form the canonical red-team specification
Paṭṭhāna Seventh book of the Abhidhamma Source of the twenty-four paccayas
Visuddhimagga Commentarial encyclopedia (Buddhaghosa, 5th c.) Source of the near-enemies catalogue + three-depth-defilement articulation

The vocabulary is large; the institutional summary uses the load-bearing subset. The companion research paper (Abhidhamma-Layer Implementation Mechanisms) develops each term's engineering interpretation in full.

2 · What the position implies for HeartBank's alignment work

Three operational commitments follow.

First, HeartBank's alignment-research investment is weighted toward the cognitive-mechanism layer rather than toward marginal improvements at the training-method layer. The training-method layer is mature; the cognitive-mechanism layer is where unsolved alignment problems live, and HeartBank invests accordingly.

Second, HeartBank treats interpretability-as-subtraction — the targeted dissolution of misalignment-generating circuits — as a methodologically more substrate-native posture than preference-learning-as-augmentation. The wholesome-roots of the Abhidhamma's analysis are apophatic (alobha, adosa, amoha — non-greed, non-hate, non-delusion); alignment is the absence of distortion, not the addition of value-encoding. This shapes which research directions HeartBank supports.

Third, HeartBank's red-teaming and evaluation work uses the structural-mimicry frame — every alignment target generates a characteristic near-enemy, and finding the near-enemy is first-class safety work. The brahmavihāra near-enemies (the Visuddhimagga's catalogue) are the canonical example of this discipline; HeartBank's evaluation extends the catalogue to contemporary alignment targets (helpfulness ↔ sycophancy; honesty ↔ pedantic literalism; harmlessness ↔ vacuity; etc.).

3 · Stance toward the field

HeartBank engages the contemporary alignment community as a peer institution contributing a complementary vocabulary, not as an outsider importing a foreign framework. The cognitive-mechanism-layer work the field is now attempting is good and serious work; the Abhidhamma supplies it with a vocabulary the field is otherwise composing in pieces. HeartBank's position is that the two efforts converge — and that the convergence is structurally substantive, not coincidental.

Honest limits, stated before the field states them for us

Nothing has been demonstrated. The institution has produced no interpretability result, has not expressed any known mechanistic finding in Abhidhamma vocabulary, and has not shown that doing so makes a single problem more tractable. This is a claim about a vocabulary's suitability made by a party that has not yet used it to do the work.

The instructive precedent is serious and it did not succeed at this. The Embodied Mind (Varela, Thompson and Rosch, 1991) is the most rigorous previous attempt to bring Buddhist phenomenological analysis into cognitive science. It was influential, it is still read, and it did not become the working vocabulary of the field it addressed. Conceptual import is adopted only where it does work practitioners cannot do otherwise; being right about a map's aptness is not sufficient.

The honest test is therefore a deletion test: does the term do work no existing interpretability term does? Features, circuits and superposition were derived from the artifacts rather than imported, and they are what the field is building with. Cetasika-style decomposition of a mental event into co-arising factors is a candidate precisely because the field lacks a settled account of how simultaneous features compose — but the burden is on us to show it on a real model. ⚠️ If the mechanism survives with every Pāli term removed and the term is doing no work, the critics are right.

What would change this: if interpretability arrives at a settled compositional account using its own terms, and that account handles the deceptive-alignment and near-enemy cases this position says the surface layer cannot reach, the Abhidhamma vocabulary is an elegant redundancy and the correct response is to drop the claim. The institution would rather that happen than be right.

4 · References