What this concept is actually saying
The canonical entry says it in one line: 聲 (śabda) is the auditory sense-object that serves as the object-field of ear-consciousness (耳識) — one of the six dusts (六塵, lit. "six kinds of dust") that pair with the six sense-organs (六根, "six sense-faculties") to make up the twelve sense-fields (十二處) and the eighteen realms (十八界). Every text in our tradition, from the *Saṃyuktāgama* 雜阿含經 to *Madhyamaka-śāstra* 中論, places 聲 as the second of the six external objects (六外入處), right after 色 (visible form) and before 香, 味, 觸, 法.
That single sentence, however, conceals a remarkable fact: 聲 is the most thoroughly dissected of all six sense-objects in the Buddhist canon. There is more about sound, from more angles, in more sections of the *Tripiṭaka*, than about any of its sister dusts. This is not accidental. Sound has an odd character: it is invisible, it enters the body without being ingested, it vanishes the moment it arrives, and it arouses emotion more directly than any other sense-object. The tradition's intense interest in 聲 reflects how dangerous — and how liberating to see through — a thing it is.
The single sentence, expanded
聲 is best understood first as a position in a four-way classification system that appears already in the early *Saṃyuktāgama* and then ramifies through every later school:
1. 聲 as 六外入處 — a neutral external field. This is its most colorless definition. 雜阿含經 says plainly: 「有六外入處。云何為六?謂色是外入處,聲、香、味、觸、法是外入處,是名六外入處。」 Here 聲 is on the same shelf as 色, 香, 味, 觸, 法: it is simply the data-point that the ear receives. In this framing, sound is morally and spiritually neutral — it is just *what is there to be heard*.
2. 聲 as 六覆 — a covering that hides the mind. The same *Saṃyuktāgama*, however, almost immediately reframes the very same object: 「色有漏、是取,心覆藏;聲、香、味、觸、法有漏、是取,心覆藏。是名六覆。」 The dust has become a cover (覆) — leaky (有漏), graspable (取), capable of *veiling* the heart. 聲 has not changed; what has changed is the relational frame: 聲 is no longer simply "there to be heard" but "there to be *taken in*, and by being taken in, it covers over what would otherwise be known."
3. 聲 as 六憂行 — a grief-course. The same text again: 「若眼見色憂,於彼色處行。耳聲、鼻香、舌味、身觸、意識法憂,於彼法處行。」 Here 聲 is fused with 耳 into the compound term 耳聲 — the root-object pair — and the compound becomes the second of the six courses of grief/sorrow (六憂行). The grammar is precise: not "the ear" alone, and not "sound" alone, but ear-sound together as a meeting that produces 憂 (worry, sorrow, agitation). This is a deeply practical observation: sound-in-itself is not what agitates; it is the ear-sound contact (觸) that triggers the agitation.
4. 聲 as 大海濤波 — a wave that can drown. Still in the *Saṃyuktāgama*, the same object becomes a wave: 「耳、鼻、舌、身、意是人大海,聲、香、味、觸、法為濤波,若堪忍彼法濤波,得度於意海。」 Note the precise reversal: the sense-organs are now the ocean (大海), and the objects are now the waves (濤波). The "self" doesn't live in any one sense-faculty; it lives in the broader ocean of the six senses, and sound is one of the waves that can roll over that sea.
These four frames — neutral field, covering, grief-course, drowning wave — are not four different concepts. They are four descriptions of the same object seen from four different angles of contact (觸). Once 聲 makes contact with 耳, it can be a door, a covering, a sorrow, a wave; which it becomes depends on what the mind does at the point of contact.
Then the *Ekottarāgama* turns the lens again: 聲 as 五欲
The five desires (五欲) of *Ekottarāgama* 增壹阿含經 list 聲 as the second of the five: 色, 聲, 香, 味, 細滑. The text stages a famous scene: five kings each nominate one of the desires as supreme, and ask the Buddha to settle the matter. The Buddha refuses to crown any one:
「若復有人性行著聲,彼聞聲已,極懷歡喜而無厭足,此人於聲最妙、最上,五欲之中聲最為妙。」
The Buddha's verdict is not that 聲 is or is not the best desire. It is that whether 聲 is the most wonderful thing depends on the 性行 (natural inclination, temperament) of the hearer. To someone constitutionally drawn to sound, sound is the highest and most wonderful of the five. Crucially, the text continues:
「若復彼人性行著聲,爾時彼人不著色、香、味、細滑之法。」
The five desires are mutually exclusive. Being attached to sound *prevents* attachment to color, fragrance, taste, and touch — and vice versa. This is not a hierarchy; it is a zero-sum attentional economy. The Buddha is not giving an answer to the kings; he is showing them that the very structure of *asking which desire is best* presupposes a self already captured by欲望 (desire). See the frame, he is saying, and the question dissolves.
This "mutual exclusion" teaching is profoundly counter-intuitive for the contemporary reader, who usually assumes that the more sensations, the worse; the senses are imagined as cumulative stressors. The *Ekottarāgama* is saying something subtler: the senses are mutually substitutive. Replace one cage with another and you have not been freed.
Then the *Vajracchedikā* lifts the question entirely
The Diamond Sūtra 金剛般若波羅蜜經 doesn't multiply the angles — it strips them:
「若以色見我,以音聲求我,是人行邪道,不能見如來。」
Note the term-form: the Kumārajīva translation uses 「音聲」 (the two-character compound *sound*, *vocal sound*, *audible expression*) rather than the bare 「聲」. This is a stylistic feature of this specific translation. More importantly, the Buddha's line refuses the entire question: whether you seek me through visual form (色) or through vocal sound (音聲), you are on the wrong path. The Tathāgata is not the sum of perceptible qualities; therefore not 色, not 聲, not any compound of perceptual objects. This is not, as some later readers have mistakenly believed, a condemnation of mantra, chanting, or vocal practice. The point is ontological: don't mistake the audio-image for the reality.
The *Mahāratnakūṭa* 寶積部 — both the boundless and the echo
The ***Mahāratnakūṭa-sūtra* 大寶積經 holds 聲 in two hands at once. On one side, the Tathāgata's voice is described as boundless (無量)**:
「猶如虛空普周無邊……如來音響不可限量。」
This is 聲 at its maximal extension — a sound that is coextensive with space, that does not diminish with distance, that has no edge. On the other side, the same text gives the line that establishes the entire 觀聲如谷響 (contemplate sound as valley-echo) tradition:
「觀聲如谷響,其性不可得,諸法亦如是,無相無差別,了知皆寂靜,是名聲三昧。」
This is also the locus where the named practice 聲三昧 (Sound Samādhi) appears — note that this term belongs specifically to the 大寶積經 vocabulary and is not found in the other canonical sources. It establishes a doorway: through contemplating sound as a valley-echo, one enters the same emptiness as all phenomena. Sound becomes a *gateless gate*.
The same text repeats the metaphor: 「猶如深谷聲,其響無有實,是故不著世,如是觀世間。」 「彼聲無有實,而於中聽聞,人尊宣說此,救拔諸凡愚。」
The pairing is exact and must not be collapsed: 音聲無量 is sound as the Tathāgata's skillful function; 聲如谷響 is sound as it appears to ordinary perception. Holding these two together prevents two opposite errors — on one side, mistaking any particular audible voice for the Tathāgata (the Vajracchedikā's warning); on the other, dismissing all vocal Dharma as empty (a misreading that would orphan the entire oral and liturgical tradition).
The *Śūraṅgama-sūtra* 首楞嚴經 — the most surgical cut
The *Śūraṅgama-sūtra* handles 聲 with the most clinical precision of any text. It does something the early āgamas do not: it separates two words that are usually fused in speech — 「聞」 (hearing, the ear-faculty's inherent capacity-to-hear) and 「聲」 (the external sound-dust that meets the ear). The two are not synonyms. They are two different categories of thing.
The Buddha sets the two terms against each other through the famous bell-striking parable. The disciple Ānanda answers: when the bell is struck, there is sound and hearing; when the sound dies away, both sound and hearing are gone. The Buddha demolishes this:
「聲銷無響,汝說無聞;若實無聞,聞性已滅同于枯木,鍾聲更擊,汝云何知?」
If the hearing-nature (聞性) truly disappeared with the sound, you would be a dried piece of wood. Yet you can hear the *next* bell strike. Therefore the hearing-nature has never come and gone along with any sound. The text is explicit:
「聲於聞中自有生滅;非為汝聞聲生、聲滅,令汝聞性為有、為無。」
In other words: sound comes and goes inside the hearing-nature; the hearing-nature does not come and go inside the sound. This is one of the most important distinctions in the entire Buddhist corpus on perception, because it shows that the capacity-to-hear is not a product of any particular sound. It is prior to all sound. Mistaking the two is called in the *Śūraṅgama* the inverted view (顛倒):
「汝尚顛倒,惑聲為聞……」
To "confuse sound for hearing" is the foundational perceptual error: assuming that one's auditory awareness *is constituted by* the sounds it processes, rather than recognizing that the awareness is the unconditioned, and the sounds are conditioned.
The text then frames the practical instruction in terms of *abandonment (棄)*, not annihilation. Sound belongs to the side of arising-and-ceasing (生滅邊): 「不循所常,逐諸生滅,由是生生雜染流轉。若棄生滅,守於真常,常光現前,塵根、識心應時銷落。」 To abandon the side of arising-and-ceasing is to guard the always-so (守於真常). This is not anti-sound; it is the procedure of stopping confusing the conditioned wave for the unconditioned ocean.
The *Abhidharmakośa* 俱舍論 — eight kinds, no karmic seed, no particles
The ***Abhidharmakośa* 阿毘達磨俱舍論 gives the phenomenological anatomy** of 聲:
「色二或二十,聲唯有八種。」
Of the twelve āyatana, only eight kinds of sound are discriminated. (In the Sarvāstivāda phonetic analysis this relates to articulated vocalizations — voiced/unvoiced, aspirated/unaspirated, etc., but the precise eightfold list varies across commentaries; the point here is that sound is technically classifiable in a way that color, taste, and touch are not — there are fewer irreducibly different kinds of sound than of color.)
Three technical features that are crucial and that contemporary readers should not skip:
1. 聲 is non-karmic in origin (無異熟生) — it is not directly produced by past karmic ripening; it must wait for the conjunction of conditions. Compare this with the five inner sense-organs, which are both *matured by karma* and *sustainably nourished*; sound has neither property. Its arising is more contingent than the arising of seeing or hearing itself. 2. 聲 is non-obstructive (無礙) — unlike visible form, sound does not physically block other sound. Two sound-waves pass through each other. This marks 聲 as belonging to a different ontological category from 色. 3. 聲 cannot be reduced to particles (極微): 「欲微聚無聲。」 In the smallest conceivable material unit, there is no sound. Sound appears only at the level of assembly. This is a structural feature that will matter when we place 聲 next to 色 in our contemporary walkthrough.
The most radical move in the *Kośa* is its discussion of 名 (name), 句 (sentence), 文 (syllable/character) — whether the linguistic unit is or is not identical to sound. After many pages of analysis, the conclusion is:
「聲即是名,此名安布差別為頌。」
Sound is the substrate of name; it is not itself separate from name. The arrangement of sound-differences into a poem (頌) is the formation of 名身, but no separate "name-thing" exists over and above the sound. This is significant for our contemporary meditation on the page: there is, in the strict abhidharma sense, no abstract linguistic entity floating behind the voice. When you say "I love you," the love is not in the abstract sentence-as-object — it is in the sound-arrangement-as-event, the vocal act, and the listener's response.
The *Mahāprajñāpāramitā-śāstra* 大智度論 — who hears, anyway
The ***Da zhidu lun* 大智度論 asks the question contemporary neuroscience often poses as if it were novel: who hears the sound?** It then shows that the question is unanswerable the moment it is taken literally.
Three "candidates" are tried in turn: 1. If the ear-organ (耳根) hears, the ear-organ has no cognition of its own — it cannot smell or taste; how could it hear? 2. If ear-consciousness (耳識) hears, ear-consciousness is one-moment and does not *discriminate* (分別) — it cannot weigh or compare, so how could it hear as we understand hearing? 3. If mind-consciousness (意識) hears, mind-consciousness does not directly perceive present objects — the blind do not lose their mind but cannot perceive sound.
The conclusion is sharp and should not be misread as mysticism:
情、塵、意和合 — ear (情) + dust-object (塵) + intention (意) — then ear-consciousness arises; only then does mind-consciousness follow and discriminate. No single one of the three hears alone.
This is not exotic metaphysics. It is the experiential observation that hearing requires (a) a physical apparatus, (b) sound-waves reaching it, (c) a directed attention. Remove any one and there is no hearing. Modern auditory neuroscience would, with different vocabulary, agree with the three-fold structure, although the specific interpretation differs — and crucially, **the *Da zhidu lun* does not claim that science proves its point**, nor that its metaphysical analysis is identical to neuroscience's signal-processing account. The compatibility is structural; the accounts remain different in kind.
The valley-echo (谷響) is then extended to a physiological example: the *uvula* (憂陀那) sends breath back to the navel, and the navel produces a resonant response — speech is assembled from the resonance of many contacts, like an echo in a closed chamber:
「一切聲皆是眾緣和合之假相,無有實作者。」
The summarizing verse, used across East Asian Buddhism, ties it all together:
「觀聲如呼響,身行如鏡像;如此得觀人,云何而不忍?」
Look at sound as a calling-back-echo; look at bodily action as a mirror-image. Once the seeing is so, how could one *not* feel compassion?
The *Yogācārabhūmi* 瑜伽師地論 — the moment, the flash, the holy words
The ***Yogācārabhūmi-śāstra* 瑜伽師地論 expands the analysis along a different axis: classification and momentariness**.
The 「or立一種乃至或立十種」 — *one way of classifying up to ten ways* — gives an exhaustive list that is itself a meditation on the irreducibility of sound-as-phenomenon:
- One kind: by virtue of being what the ear processes. - Two kinds: definitive (了義) and non-definitive (不了義) speech. (This carries forward into later commentarial use of 「了義」 and 「不了義」 for sutra types.) - Three kinds: sound from a felt-bodied-element (e.g., human voice — internal), from an unfelt-bodied-element (e.g., wind in a pipe — external), and from both. - Four kinds: wholesome, unwholesome, covered-indeterminate, uncovered-indeterminate. Or: four holy words (四聖言聲) — "seen is seen, heard is heard, felt is felt, known is known" — and four un-holy words — their negations. - Five kinds: by the five realms of rebirth. - Six kinds: recitation-and-memorization, inquiry, teaching, dialectic, confession, and noise. - Seven kinds: male, female, low, middle, high, bird/beast, wind/forest. - Eight kinds: the four holy and four un-holy again. - Nine kinds: by temporal and spatial position. - Ten kinds: the five-musical-entertainment sounds (五樂所攝聲) — drums, strings, singing, dancing, female and male performers.
Notice what this list reveals: 聲 is never abstract. Every classification circles back to a concrete instance of speech, music, noise, or expression. This is a discipline lesson for students of the text: the Buddhist analysis of sound always lands back in the lived, audible world.
The single most striking feature in this sūtra is the description of 聲's momentariness (剎那生滅):
「諸聲纔宣發已,尋即斷滅,故於色聚中不恆相續。」 「隨所聞處遍滿頓起,如焰光明,非漸漸生展轉往趣。」
Every sound arises in a flash, fills the space where it is heard, and ceases — like the flash of a flame, not like a wave that travels. This is a structural observation that contradicts the everyday intuition that sound "travels". The text is right in a subtle sense: what we think of as a sound-wave traveling is, at the percipient end, a series of discrete momentary collapses into hearing. The "travel" is a conceptual overlay.
It also articulates the version of the "I-through-voice" error:
「若以色量我,以音聲尋我;欲貪所執持,彼不能知我。」
If you try to measure me by visible form, or to track me down by sound, you are held tight in the grip of desire — you will not know me.
The *Mūlamadhyamaka-kārikā* 中論 — the dependent-arising of sound
The 《中論》 places 聲 into the dependent-origination framework by reference to the prior analysis of seeing:
「耳鼻舌身意,聲及聞者等,當知如是義,皆同於上說。」
"Catch the meaning of all these [ear, nose, tongue, body, mind; sound, smell, taste, touch, thought, and the hearer, etc.] by analogy with what has been said above." The 中論 refuses to analyze 聲 *in itself*; it dissolves 聲 by reminding the reader that the same logic that applies to eye-color-seer applies to ear-sound-hearer. Then it adds a specific argument: if the seer were the hearer (and the feeler), they would be one single self (神); if so, the eye could hear sound. But this is not how things are. Therefore no such self exists at the basis of the senses.
This is the 中論's elegant way of saying: even if you find your "self" through sound, the self you find is the same self-as-seer you would otherwise find through vision; but no such single self-entity is locatable in either. Sound as a doorway to self-discovery is, for the 中論, just another door that opens onto the same empty hall.
The map, condensed
Reading across all nine principal sources we have four doctrinal anchors:
1. 聲 is a sense-object that becomes, at contact, a covering, a grief-course, a wave, a desire, an empty echo. The transformation is not in the object; it is in the *frame of contact*. 2. The 聞性 vs 聲塵 distinction (Śūraṅgama) is the most important single cut. The capacity-to-hear is not constituted by any sound. The mistake of mistaking sound for hearing is the root perceptual inversion. 3. 聲 is momentariness-without-particles — it cannot be reduced, it cannot be located, it cannot persist. It is the most clearly impermanent of the six dusts. 4. There is no "I" discoverable through sound — whether as Tathāgata (Vajracchedikā), as conditioned self (大智度論), as one-mind-across-senses (中論), or as bounded ego (瑜伽論). All four routes arrive at the same conclusion: 聲 is an excellent negative mirror.
Where 聲 sits in the path of liberation
If you hold this matrix and ask *where exactly 聲 functions in the 12-link dependent arising 十二因緣*, you see that 聲 enters at the very junction where 觸 (contact) meets 受 (feeling): 六入 (sense-fields) → 觸 (contact) → 受 (feeling) → 愛 (craving) → 取 (clinging). The whole machinery of suffering pivots on what happens at the moment ear meets 聲. Compared to 色, 聲 has a special edge: it vanishes while still triggering affect. You can be wounded by a word that has already completely ceased as sound. This is why the tradition devotes so much literature to it — and why the *Sūtra of the Wise and the Fool* family and the Vinaya devote pages to the vocal karmic acts (speech lies, gossip, harsh words, idle chatter), giving sound a place in the ethical architecture parallel to bodily action.
Liberation from 聲 (as object of desire) does not mean the abolition of hearing. The early Buddhist goal is 非我也、非我所 — not mine, not my self. When 聲 arises, it arises; when it ceases, it ceases. The covering drops when the hearer stops confusing themselves with what is heard.
Life walkthrough — putting the classical framework on a contemporary day
Let us take the day as an audio-archeologist might: closing the eyes, walking through the soundscape, naming each one in classical terms. I will use five scenes. As we go, note where the contemporary mind most often *mis-maps* the classical term — those mis-mappings are where the teaching wants to land.
Scene 1 — The morning alarm and the audio commute
What is happening. Six-thirty-something. The phone makes a sound: that single ping of the alarm. Sleep shatters. The body jolts. Within seconds, the air fills with a podcast episode, a Spotify playlist, news headlines — a stream of voices and music that will last the entire commute.
Mapping by classical category. - The alarm is 六入處中的聲 — neutral external field, *if* one experiences it only as data. But the moment the body jolts, it has already become 六覆 (six coverings): the alarm covers the dream-state of sleep; it covers the slow emergence into the day. Sound here acts as a covering. - The podcast-over-shoulder-bag that lasts forty minutes is 六憂行中的耳聲: a continuing grief-course, perhaps low-grade (the news, the colleague's complaint). The compound "耳聲" is exactly right here — *not* the ear alone, *not* the podcast alone, but the *meeting*, in the car, of an exhausted person and a stream of content designed to keep them occupied. - The playlist is 五欲之聲 (the second of the five desires): chosen precisely because of *性行* — to a constitutional music-lover, music is the highest and most wonderful; to a constitutional silence-lover, music is the highest torment. The same playlist means opposite things on Tuesday and Saturday depending on 性行.
Where the mis-mapping happens. The contemporary reader hears 「六覆」 and assumes the teaching is *just turn off the alarms, just quit the podcast*. The sūtra does not say this. The *Da zhidu lun* triple analysis (情+塵+意) is the precise correction: what makes the alarm a covering is the conjunction of (1) an organ rested enough to wake, (2) a sound sharp enough to startle, (3) an intention that has given the alarm permission to startle. Turn off only (3) — give the alarm no meaning, let the body wake on its own — and the same acoustic event ceases to be a covering. It still sounds; it just no longer covers.
This is the place where the *Śūraṅgama* line 「聲於聞中自有生滅;非為汝聞聲生、聲滅,令汝聞性為有、為無」 becomes experientially relevant. The alarm's *sound* comes and goes inside the hearing-nature, which is the same nature that was present in deep sleep, in the dream, in the moment of jolt. The alarm did not create hearing. The alarm did not create *you*. The classical frame liberates the moment from being a story about alarm-sound to being a story about the hearing that contains alarm-sound.
Scene 2 — The argument with the person you love
What is happening. A word. "You always…" The voice rises. The sentence completes — already vanished as sound by the time the next sentence starts. But the after-image keeps running.
Mapping by classical category. - The word as 六覆 has acted twice: it has covered what the speaker might have actually needed to say; and it has covered the listener's capacity to *hear* the underlying need. Both coverings occur because the word-as-聲 *replaces* the seeing-of-the-person. - The compound 耳聲 in 六憂行 is precisely the right classical frame here, not 聲 alone. The argument is not in the sound-as-such; it is in the contact: the lover's ear meeting the lover's voice meeting a hurt (情、塵、意). Without 意 (the directed intention of engagement), the same waveform from the same mouth would be only - to use the *Śūraṅgama* distinction - a *聲塵*, not yet an *耳聲*. The transformation from data to grief happens at *meeting*. - The 瑜伽論 warning: in the ten kinds of sound, the four *un-holy* (非聖言) include the deliberately obscuring word and the harsh word. 「四非聖言聲: 顛倒虛誑之語」 — the *Yogācārabhūmi* text itself calls out a speech-act that *inverts reality* as one of the *four un-holy sounds*. This is not a generic ethics-of-speech list; it is a *sound-classification*. The argument-the-sentence is, in this schema, a teaching moment in real time about exactly which category one's voice is entering.
Where the mis-mapping happens. A contemporary reader, working through mindfulness traditions, often thinks 「六憂行」 means "avoid the sound." The classical instruction is different. The *Saṃyuktāgama* itself frames 六憂行 *inside* the project of "堪忍彼法濤波" — **enduring the wave *as* a wave, in order to cross the ocean. The teaching is not "make the wave go away" but "remain in the wave without being rolled by it." The 大海濤波 metaphor inverts the apparent metaphor: the senses are the ocean, the dusts are the waves, the waves are what you must learn to *swim in*, not escape**.
So: not "stop arguing," but "stop confusing the *sound-that-ceases* with the *hearing-that-does-not-cease* (聞性)." When the harsh word lands, classically, it lands in ear-sound contact (耳聲) and triggers 憂. The next move, classically, is to notice: *the sound is already gone*. What remains is not sound but 識-with-form-of-sound (想蘊 + 行蘊): the mental echo of the spoken word, which the *Abhidharma* classifies as a *意* (mind-object) event rather than an *耳* event. The work is to localize: *what is being reacted to is now happening in 意, not 耳*. That localization is the immediate attention-shift the *Śūraṅgama* names 「捨生滅, 守真常」 — abandoning the side of arising-and-ceasing, guarding the always-so.
Scene 3 — Voice messages, voice cloning, and the literalization of 谷響
What is happening. A colleague sends a voice message you cannot tell apart from their actual voice. A podcast plays an interview segment that turns out to be a synthetic voice. The cloning tool has reached the point where the original speaker cannot reliably tell the difference.
Mapping by classical category. - The *Mahāprajñāpāramitā-śāstra*'s 谷響 metaphor is being literalized in contemporary technology. 「響事空, 能誑耳根」 — *the echo-event is empty, it can deceive the ear*. The first-generation deception was geographical: an echo in a mountain valley that sounded human but was no human. The contemporary form is engineered: a *synthetic* voice that the ear cannot tell apart from a *natural* one. - The *Abhidharmakośa*'s final verdict — 「聲即是名,此名安布差別為頌」, **sound *is* name**; there is no separate abstract "name-thing" — becomes an entirely practical matter. The synthetic voice *is* the message-as-such; there is no original "speaker" behind the engineered voice waiting to be found. The voice-pointing-mistakenly-to-author reaches a hypertextual version of 瑜伽論's warning: 「若以色量我,以音聲尋我……彼不能知我」 — you cannot find me through sound. - The *Da zhidu lun* analysis becomes operative: who hears the message? At the level of neural processing the answer is: a vast signal-processing pipeline trained on millions of voices which disassembles the waveform into phonetic features, words, semantic content, prosodic affect. At the level of Buddhist analysis: 耳根 + 聲 + 意三緣和合, then 耳識起, then 意識分別. The synthesis of voice collapses the assumption of a unitary "speaker"; the analysis of hearing dissolves the assumption of a unitary "hearer."
Where the mis-mapping happens. The contemporary reading often reverses the *Da zhidu lun*'s point and says, "well, the speaker is just the algorithm." This is the 神 (self) trap. The Buddhist analysis is not that the algorithm *is* the speaker; it is that there is no speakers-at-all locatable in the voice-event as such. The voice + the algorithm + the listening ear + the trained model of "human voice" + the intent of the listener all conjoin. To say "the algorithm heard" is the same error as "the ear heard" — it's a partial heap mistaken for a self. What is being literalized in the AI voice moment is, classically, the teaching of the mutual arising (緣起) of voice, speaker, hearer, and meaning.
Note carefully: this is *not* a claim that AI voice-cloning is a Buddhist revelation or that the science "proves" the doctrine. The *categories* in the classical texts are structural; they predate and outlast any specific technology. The compatibility is instructive; the identification is not.
Scene 4 — The sound that loops after it's gone
What is happening. A song from five years ago. A voice from someone who has died. A critical sentence from a parent. The sound itself is already over — but the loop in the mind is not over.
Mapping by classical category. - The *Yogācārabhūmi* line 「諸聲纔宣發已, 尋即斷滅」 is the single most useful sentence in the entire canon of 聲 for this case. The sound has already ceased. It is literally, ontologically, gone. There is no remaining auditory event. - What is looping is therefore not 聲. What is looping is 識 (consciousness) recombined with 想 (perception: 「取相識別」, the *recognition-by-taking-on-features*) and 行 (volition/construction) — the *Āgama*'s definition of the fourth and fifth aggregates: 想 is **the marking function that *recognizes* this as "that voice", 行 is the karmic-affective momentum that keeps replaying it. What hurts is not sound; what hurts is the *construction* of the sound-as-thing-in-mind.** - This is the precise place where the *Śūraṅgama*'s 「惑聲為聞」 inversion does its therapeutic work. The mind has confused the sound for the hearing, but it has done something subtler too: it has confused a sound-event for a permanent object. The trauma-loop is the most common contemporary form of the 八顛倒 (eight inversions) involving sound. To recognize, classically, that 「聲於聞中自有生滅」 is to begin to separate the *event* from the *hearing*, and the *hearing* from the *constructed-image-of-hearing*.
Where the mis-mapping happens. Trauma research and Buddhist practice share terrain here, but they should not be conflated. Modern trauma research observes that flashbacks occur via reconsolidation of memory; the *Āgama* and the *Yogācārabhūmi* observe that what is being reconsolidated is already a mind-object, not an ear-object. The contemporary neuroscientist says "the amygdala reactivates the memory network"; the classical Buddhist diagnostic says "the 識-with-想 contact-loop is feeding on a 聲-as-marker that is past." Both point to the same finding by very different vocabularies. The practical instruction in the Buddhist tradition is not to suppress the loop (which strengthens it) but to localize: *what is occurring is not sound; what is occurring is mind-construction around a sound that no longer exists.* When this is seen directly, the loop releases. (Modern therapies such as EMDR, IFS, and somatic-processing operate by similar localizations through different vocabularies; here again, *compatibility* is not *identity*; classical practice is its own practice with its own yardsticks.)
Scene 5 — The lulling voice and the directed speech
What is happening. A bedtime story. A teacher. A recorded meditation instruction. An act of confession between friends.
Mapping by classical category. - The four holy words (四聖言聲) from the *Yogācārabhūmi*: 「見言見, 不見言不見, 聞言聞, 不聞言不聞, 覺言覺, 不覺言不覺, 知言知, 不知言不知」 — *what is seen, called seen; what is not seen, called not seen; etc.* This is not just ethical advice; it is the defining feature of sound-as-such-holy. The holy word is the one that performs no transformation of the real. The parent saying "I see you" when the child is unseen; the teacher saying "I do not yet know" rather than faking knowledge; the friend saying "this hurt me" rather than performing indifference — each is a holy sound. The 大寶積經's line 「彼聲無有實, 而於中聽聞, 人尊宣說此, 救拔諸凡愚」 — though that sound has no essence and yet is heard, the Holy One teaches this to rescue the foolish — applies here: the rescue comes not from the "essence" of the sound but from its function. - Conversely the lullaby-as-flattening (the soothing voice used to *prevent* the child from expressing distress rather than to hold the space for it) is the opposite of the holy word: it is a masking sound (覆), an inverting sound (顛倒), one of the four un-holy words (四非聖言聲) in real-time pedagogy. - This is also where the *Śūraṅgama*'s 「常光現前」 has its collective effect: when a community or a family treats the 見言見 discipline as the floor, the room itself becomes an environment where 聞性 can stabilize — the children, the friends, the partners all begin to construct a relational hearing-nature that isn't thrown around by every wave.
Where the mis-mapping happens. Many readers of Buddhist speech-ethics treat the 四聖言 as a list of prohibited and permitted speech-acts (don't lie, don't gossip, etc.). This is a flattening. The *Yogācārabhūmi*'s emphasis is on the manner more than the content: a true statement said with an intent to dominate is not a 聖言; a difficult statement said with the intent to disclose truth is. The classification has to do with the directionality of the mind behind the sound, not the categorical character of the words uttered.
Why contemporary people need this teaching
The contemporary situation has produced, for the first time in human history, a near-total saturation of 聲. Three structural shifts make the classical frame relevant in new ways.
1. Acoustic saturation has outpaced any prior human condition. For most of human history, the moment of waking silence was a given — even farmers and workers had pockets of silence. The contemporary urban condition, *especially* the earbud-in condition, produces a nearly continuous audio-strip from morning alarm to sleep-podcast. The classical *Saṃyuktāgama* categories — 六覆 as a real-time diagnostic, 六憂行 as a continuous-grain identification — are now describing not rare moments but the modal experience. Practically speaking: most contemporary 憂 (sorrow, agitation) is now streamed rather than episodic, and the audio-stream is the most under-recognized contributor. **What the Buddhist tradition names as 聞性 (hearing-nature) is, for contemporary people, mostly *not yet known as their own*; it is being permanently overlaid by a stream of 聲 that they have not learned to recognize as such.**
The instruction 「守於真常」 has a contemporary sound to it that the ancients did not need: it means to recognize the hearing that is there before, during, and after any sound — and to *let that recognition stabilize*. This is not the same as wearing noise-cancelling headphones or scheduling "quiet time" (though those may help). It is the recovery of the hearing-capacity as one's *own*, not as one's *contents*.
2. The collapse of the speaker/instrument distinction. Until recently, the assumption that a sound had a *real speaker in space* was roughly correct. The contemporary voice-cloning, voice-synthesis, and audio-deepfake environment has rearranged this. The 大智度論's three-way analysis (情+塵+意) was already a deconstructive reading of "who hears"; the *contemporary* situation extends the question to "who speaks?". When the answer *cannot be determined from the waveform*, the classical 谷響 teaching stops being a metaphor and becomes the only descriptive reality. Practically, this means a contemporary person who wants to relate to audio-content ethically *must* develop the 大智度論-style deflation of the speaker-claim: not "the algorithm is the speaker," not "no one is the speaker," but "the speaking-as-event conjoins condition, and 'speaker' is a name for one of the conjoining conditions, not for an entity behind them."
3. The trauma-loop problem has new cultural reach. The contemporary psychological literature on trauma-loop, intrusive recurrence, and rumination describes a population-level phenomenon that the Buddhist *Āgama* and *Yogācārabhūmi* had specific things to say about. The *Saṃyuktāgama*'s 「諸聲纔宣發已, 尋即斷滅」 applied to a single monk or nun in a single moment of meditation is one scale; the same passage applied to a population saturated in audio-headphones whose audio-streams are interrupted by alerts whose audio-streams pile up into reveries whose mental echoes haunt an entire culture is another. What was once a meditation practice for monastics is becoming *a meditative hygiene for the audio-saturated*. The classical diagnosis — *what you are looping is not sound* — is also a contemporary practical, not only doctrinal, instruction.
These three reasons do not constitute a proof that *the Buddhist view is correct*; they constitute a description of why the structure of the classical analysis maps well onto the present moment. The classical diagnosis of 聲 will continue to be irrelevant where the actual audio-condition of a person does not implicate the categories — for someone in a remote agricultural setting, most of this would be over-engineered. For someone in a contemporary audio-saturated environment, it is under-engineered if not addressed.
Common misreadings and the corrections that the canonical sources themselves offer
**Misreading 1. 聲 means *speech*, and therefore the analysis of 聲 is an analysis of *speech-ethics* alone. The 瑜伽論十種 sound classification shows the opposite: 聲 includes wind-in-forest-sound, animal-sound, drum-sound, music-sound, machine-sound. 聲 is the entire audible range, not the human voice. Speech is one of its categories (受持演說聲, 論義決擇聲, etc.). When the text talks about 五欲之聲, it means literal audible music, recorded or live**. The contemporary reader who hears the sūtras as a treatise on lying is reading only the tip.
Misreading 2. The instruction 「以音聲求我, 是人行邪道, 不能見如來」 is a condemnation of chanting, mantra, or vocal Dharma practice. The Vajracchedikā line is about *seeking the Tathāgata through sound as a defining characteristic*. It is not a decree against vocal practice. The 大寶積經's extensive treatment of the Tathāgata's own 無量音聲 (boundless vocal sound), and its entire establishing of 聲三昧 as a practice, demonstrates that within the same canon the Buddha is described using voice, extensively, as an instrument of teaching. The point of the Vajracchedikā line is ontological: the Tathāgata is not constituted by voice; therefore don't mistake any voice-as-such for the Tathāgata. This is structurally the same warning as against mistaking a particular body-as-form for the Tathāgata. Reading it as anti-mantra is collapsing a metaphysical nuance into a discipline prohibition.
**Misreading 3. 聞性 means *the ear* or *the auditory system*.** The Śūraṅgama-sūtra is at great pains to distinguish 聞 (the *hearing-nature*) from 耳 (the *ear-organ*), from 聲 (the *sound-object*), and from 耳識 (the *ear-consciousness*). 聞性 is a technical term for the inherent Buddha-nature as it manifests in the auditory modality — it is *not* the ear, *not* the nervous system, *not* the auditory cortex. The canonical move 「聲銷無響, 汝說無聞;若實無聞, 聞性已滅同于枯木」 makes this crystal clear: if 聞性 *were* the ear-as-organ, the dried-wood objection would be trivial — of course a dried ear cannot hear; that's nothing to do with 聞性. The point is that hearing is not contingent on any organ: the present capacity-to-hear is not produced by the ear. The contemporary reader needs to feel the full strangeness of this claim: the ear-as-organ is a condition, but it is not the *hearing*. What 聞性 names is the hearing prior to all sound.
Misreading 4. "Sound is bad, so we must seek silence." None of the canonical sources we have examined teaches *silence-as-the-goal*. The 大海濤波 metaphor explicitly tells the practitioner to *endure* the wave in order to *cross the ocean*. The *Da zhidu lun* triple analysis is a *description* of contact, not a *prohibition* of it. Even the *Śūraṅgama*'s strong line 「若棄生滅, 守於真常……塵根、識心應時銷落」 is a description of what happens when one abandons confusion, not when one abandons sound. Six classical sources, read carefully, agree on this: the goal is not silence, it is the unmuddling of contact. A practitioner who retreats to silence-as-policy has, classically, *not yet* noticed what the teaching is about. They have confused "no sound" with "no confusion with sound."
Misreading 5. 五欲的相互排斥 means "pick the lesser of five evils." The *Ekottarāgama* mutual-exclusion structure is subtler: **the five desires are mutually substitutive *in the same structure***. If you swap obsessive podcasting for obsessive social-media-image-collecting, you have *not* practiced the teaching. The teaching is that whichever of the five one is constitutional toward (性行), that one will be the most wonderful to them, and the others will be comparatively inert. The Buddha is not giving a relative ranking; he is pointing at the *fact of constitutional-attachment* and saying: *wherever you are locked, you are locked; switching cages is not freedom.*
Misreading 6. 聲的刹那生滅 means "live as if everything is silent." The *Yogācārabhūmi* line 「非漸次往趣, 如焰光明」 describes the actual arising and ceasing of each moment of sound; it is a phenomenological observation, not a prescriptive one. The cognitive error against which it is directed is the substance-illusion (常見) — the assumption that sounds are persisting things. To notice momentariness is not to live in dead silence; it is to live in the rhythm of arising-and-ceasing with continuous knowledge of the rhythm.
Misreading 7. The "scientific proof" inference. A careful contemporary reading might note that modern auditory neuroscience, cognitive psychology, and signal-processing theory describe structures that *resemble* the classical descriptions. This is structural analogy, not identity. The *Da zhidu lun* three-fold analysis is not a signal-processing diagram; the *Śūraṅgama* 聞性 is not an auditory-cortex model; the *Yogācārabhūmi* momentariness is not a frame-rate observation. Each tradition has its own internal coherence, its own methods of validation, and its own limits. Drawing parallels is illuminating; drawing identities is a category error — and it risks either bad science (claiming Buddhism supports neuroscience that doesn't support Buddhism) or bad Buddhism (claiming neuroscience proves Buddhism). They illuminate each other at the level of structure, not at the level of truth-content.
**Misreading 8. 「聲即是名, 此名安布差別為頌」 means *vocal sound is the only kind of language*.** The *Abhidharmakośa*'s conclusion is the reverse: **sound *is* the substrate of name**; name-as-thing does not exist except as sound-arrangement. This means that *all* language-as-such is, abhidharmically, *sound-arrangement* — and conversely, that *all* sound is potentially *language*. It dissolves the modern dichotomy between "mere noise" and "meaningful speech" — at the level of analysis, both are arrangements of sound-differences. The difference is that some arrangements *carry* semantic content (頌, the poem; 句, the sentence; 文, the syllable) and others do not. This is a deeper point than it seems: the boundary between sound and language is a matter of arrangement, not of substance. This *Abhidharma* insight becomes arresting in the contemporary moment of voice-cloning and audio-fabrication: there is no "deeper" speech-essence behind the synthesized voice; the synthesized voice, if it carries meaning, *is* speech.
Coda: walking one more round
If we walk back through the day with the classical frame in hand — alarm, commute, argument, AI-voice, lullaby — what changes?
The alarm still startles, but we can know it as 情+塵+意 conjoining, and the 聞性 remains.
The commute still has its podcast, but the 六憂行 diagnostic is on the dashboard; 耳聲 is named in real-time and so is no longer *driving* in the same way.
The argument still hurts, but we can localize: the sound has ceased, the 識-with-想 is now looping; the 聞性 is unchanged; the holy/unholy line in the *Yogācārabhūmi* can be a quiet check after the words have died.
The AI-voice still deceives the ear, but 谷響 is on its proper shelf: no real speaker waits behind the waveform, only conditions and arrangements.
The song loop still returns, but with 「諸聲剎那已滅」 in hand: what is returning is not sound, and what is returning is recognized as not-sound, and the recognition itself begins to do something therapeutic that the loop had been preventing.
And none of this requires silence. It requires the *classical attention* that is the inheritance of every one of these texts: to keep hearing, without confusion, in the field of all sounds.
This is the thread that runs from the *Saṃyuktāgama* catalog to the *Vajracchedikā* warning to the *Śūraṅgama* surgical cut to the *Da zhidu lun* three-fold analysis to the *Yogācārabhūmi* momentariness to the *Mūlamadhyamaka* frame: the same single object, examined from a dozen vantages, always discloses the same teaching — *the hearing is not made of the heard; the speaker is not made of the spoken; the wave is not the ocean; and what remains, when all coverings are seen through, is something that does not come and go with any sound at all.*