What this concept is actually saying
There is a strange and beautiful moment in the *Mahāyāna-sūtrālaṃkāra* (《大乘莊嚴經論》, T1604) where the author pauses, draws a breath, and writes: "已說語成就,次說字成就" — "Speech-Accomplishment having been explained, Word-Accomplishment is now explained." This single bridging line is doing real conceptual work. The treatise has just finished describing how a bodhisattva's *general* communicative deportment (語成就, *vāk-saṃpad*) is perfected across eight qualities — that the bodhisattva does not speak coarsely, is well-tuned, is skillful, is clear, meets the audience where it is, does not angle for gain, knows the right measure, and can go on without exhaustion. Now it zooms in.
The verse that follows is the whole blueprint: 「菩薩字成就,如前義應知,聲有六十種,是說如來事」 — "The bodhisattva's Word-Accomplishment is to be understood through the foregoing sense; there are sixty kinds of sound; these are for the purpose of proclaiming the Tathāgata's [如來] matters." Three things are packed into four lines:
1. The subject is the bodhisattva (菩薩) at the stage of fully ripened communicative power. 2. The instruction is *continuity* — "如前義應知" — meaning that 語成就 gives us the orientation, and 字成就 is its finer-grained unfolding. 3. The number is sixty, and the goal is to "speak the Tathāgata's matters" — that is, to fully convey the true Dharma (正法, *sad-dharma*) to living beings.
A crucial terminological note. The English translation "word" or "letter" hides a Sanskrit pun. *Akṣara* literally means "imperishable" and covers three faces: sound (the spoken syllable), form (the written glyph), and meaning (what is signified). The treatise is doing something unusual here: it deliberately collapses "字" into "聲" (sound). It is not interested in calligraphy, orthography, or the shape of letters on a page. It is interested in what reaches the ear of a sentient being. That is why it counts out sixty kinds of sound, not sixty kinds of character or sixty kinds of meaning. The implication is striking: in the context of teaching the Dharma, what ultimately counts is not the script on the page, but the living vibration that lands in someone's ear and transforms their mind.
The prose then unfurls sixteen of those sixty sounds as exemplars — not because the others are less real, but because the treatise says "舉隅" (giving a corner) — showing the categories so the reader can extrapolate. These sixteen are not a random list. They cluster around six qualities of perfected vocal activity:
- 清淨無失 (Purity without fault): *離不正聲* — free from incorrect sound, because the bodhisattva remembers and does not forget; *不毀呰聲* — non-denigrating sound, faithful to the meaning that has been established. - 應機得時 (Timely, meeting the audience): *應時聲* — timely sound, the teaching activity arises at every appropriate moment; *不增減聲* — neither adding nor subtracting, speaking exactly the right measure for the time; *不躁急聲* — not agitated, not rushing the words out. - 饒益有情 (Benefiting beings): *歡喜聲* — joyful, so listeners don't grow weary of hearing; *善友聲* — the voice of a true friend, accomplishing the welfare of beings; *眾生根喜聲* — pleasing to the faculties of beings, so a single expression reveals countless meanings. - 相續無斷 (Continuous, unceasing): *常流聲* — flowing without interruption. - 遍滿十方 (Pervasive): *遍一切聲* — pervading everywhere, so that those near and far can equally rely on the teaching. - 圓滿成就 (Full accomplishment): *嚴飾聲* — adorned, with manifold appearances; *滿足聲* — fulfilling, so one sound can give rise to countless sounds of Dharma; *一切種成就聲* — accomplishing all kinds, expressing both worldly dharmas and ultimate meanings through analogy.
Two more sit in their own register and deserve their own hearing. *無羞聲* — "shameless-free," meaning the voice is not propped up by the expectation of offerings, patronage, applause, or career advancement. *不怖聲* — fearless, free from the inner shame and self-consciousness that makes the throat tighten. And *隨捨聲* — often translated as a kind of generous "letting-go sound," arising because the bodhisattva is skillfully entered into all fields of knowledge (*一切明處善巧入*) and so can give freely without hoarding. These three — shamelessness, fearlessness, and free-giving — are not add-ons. They are the inner condition that makes the other thirteen possible.
The whole thing reads as a kind of acoustic phenomenology of awakening. The *Mahāyāna-sūtrālaṃkāra* is asking: what would it be like if a human voice were so completely unstuck from ego, fear, and grasping that it became, in a literal sense, a vehicle for the Dharma? The answer is not "loud" or "beautiful." The answer is the sixteen qualities just enumerated — and the sixty they stand for.
This is also why the treatise places the doctrine at the end of a graduated sequence: 語成就 first (the broad etiquette and stance of communication), then 字成就 (the specific sonic enactment). The general must exist before the specific can function. You cannot have perfected sound without perfected *intention to speak*. And the ultimate horizon is not the sound itself but its telos: "是說如來事" — the speech is for the sake of the Tathāgata's matter, the proclamation of true Dharma. Sound perfected but pointed at oneself is a beautiful instrument played in an empty room.
Walking through it in contemporary life
The bodhisattva is, in classical terms, an idealized figure. But the *quality* of their speech is not mythological — it is an analysis of what perfect communicative activity looks like, and every contemporary person who teaches, parents, counsels, podcasts, manages, partners, posts, or simply *talks with intent* is somewhere on the spectrum toward or away from this accomplishment. Let us walk through a few concrete scenes.
Scene 1: A teacher on the second-to-last week of term
Imagine a high-school teacher trying to land a difficult concept — say, exponential growth, or grief, or the difference between guilt and shame. Most teachers in this situation are doing three things at once: managing the clock, managing their own fatigue, and trying to be intelligible to thirty distinct nervous systems. Watch what happens when their voice drops into alignment with the bodhisattva's sixteen.
*離不正聲 / 不毀呰聲* — They do not flatten the concept to make themselves look smart, nor do they distort it to make the students feel smart. They say what is true. A student will sometimes remember an exact sentence from a real teacher for the rest of their life; this is partly why those sentences must be technically correct.
*應時聲* — They sense that one student is on the edge of getting it and pause for her. They do not bulldoze ahead to "cover the material." The teaching activity rises at the moment the audience can receive it.
*不增減聲* — They resist both the temptation to ramble and the pressure to compress. They give the concept exactly the time it needs. This is the one modern teachers most often violate, because both the institution (curriculum, pacing guides) and the teacher (ego, anxiety) push toward either inflation or truncation.
*不躁急聲* — They do not speed up at the end of the period to "finish." The throat is not tight. The pace is human.
*歡喜聲* — The students are not bored. There is something in the voice that makes you want to keep listening.
*眾生根喜聲* — One sentence lands differently for the math kid and the poetry kid. The teacher has crafted the example so that both faculties are engaged.
*常流聲 / 嚴飾聲 / 滿足聲* — The teaching has a continuity and a richness. A single image opens into a small field of related images. The student leaves with a constellation, not a bullet point.
*無羞聲 / 不怖聲* — And — this is often invisible — the teacher is not pitching the lesson to be praised by the principal or to be safe from the parent who complains. There is no audience behind the audience. They are simply talking.
*隨捨聲 / 善友聲 / 一切種成就聲* — Because they have read widely and entered into many fields, they can bring a metaphor from jazz, or biology, or a film the students watched last weekend. The Dharma of exponential growth is taught through analogy, just as the treatise says.
*遍一切聲* — The students who are physically present and the ones at home with the link get the same instruction. Distance is not degraded.
Notice that nothing here is supernatural. It is a high bar of *intentional, transparent, audience-tuned vocal activity*, sustained over time, oriented at benefit rather than at the speaker's image.
Scene 2: A parent at the kitchen table, 9:47 p.m.
A child has just failed a test, or broken a promise, or said something cruel to a sibling. The parent has had a long day. Now apply the sixteen.
*離不正聲 / 不毀呰聲* — They do not call the child "stupid" or "lazy." These are *incorrect sounds* in the precise sense the treatise gives: they are not faithful to the meaning, they betray the actual situation, and they corrupt memory because the child will rehearse them at 2 a.m. for years.
*不躁急聲 / 不增減聲* — They do not unload a forty-minute grievance in ninety seconds. They also do not withhold everything and say "we'll talk about it later," which is itself a kind of untruth at that moment. The right measure.
*應時聲* — Sometimes the moment for a long conversation is not 9:47 p.m. The parent senses that *now* is for a brief, anchoring sentence, and the longer talk is for tomorrow morning. The teaching activity arises at the time it can actually take root.
*無羞聲 / 不怖聲* — The parent's voice does not tighten because they are worried about being judged by the child, or by their own parents, or by the Instagram version of good parenting. There is no audience behind the audience.
*歡喜聲* — Even in correction, there is a baseline warmth. Not fake cheerfulness. The recognition that the child is a being trying to become, and that being met in this moment is itself a gift.
*善友聲 / 眾生根喜聲* — The parent picks up, from the child's face, what the child actually needs: not a lecture, but acknowledgment; not acknowledgment, but a plan; not a plan, but a hug. The voice modulates accordingly.
*常流聲 / 嚴飾聲 / 滿足聲* — A child who grows up in a household where the parent's voice has this continuity and richness internalizes something hard to name. It is the felt sense that speech can be a reliable vehicle for meaning.
*隨捨聲 / 一切種成就聲* — The parent draws on their own failures, on a story from their own childhood, on a poem they half-remember. Analogy from worldly dharmas opens ultimate meaning.
*遍一切聲* — The parent at the kitchen table is also, in a sense, present to the child they will be at 7 a.m. tomorrow, and to the teenager that child is becoming.
Scene 3: A manager giving critical feedback
This is where most contemporary professional life lives. Apply the sixteen and watch almost all of them go missing.
*不毀呰聲* — the manager names what happened, not the person. *不增減聲* — they give the feedback in proportion to the actual situation, not inflated by their mood and not deflated to avoid the meeting. *應時聲 / 不躁急聲* — they do not dump it between meetings, half-present, half-already gone. *離不正聲* — they do not soften the actual issue with corporate language until the recipient cannot tell what was said. *無羞聲 / 不怖聲* — they are not angling for the report to "make them look good" to *their* boss; they are not afraid of the report's reaction. *歡喜聲* — there is enough warmth in the room that the report can actually hear. *眾生根喜聲* — they notice whether this person needs structure, story, or silence, and adjust. *常流聲 / 嚴飾聲 / 滿足聲* — the feedback is part of an ongoing relationship, not a one-off; it has texture. *善友聲 / 隨捨聲 / 一切種成就聲* — they use a metaphor from the report's own world, not from theirs. *遍一切聲* — and the values they speak in the room match the values they speak in the hallway.
The trap to spot
The single most common mapping error is to read the sixteen qualities as a checklist for "sounding good." *離不正聲* becomes "speak correctly." *歡喜聲* becomes "sound cheerful." *嚴飾聲* becomes "use rhetorical polish." This is the flattening the warning forbids. The qualities are not stylistic. They are diagnostic of what the voice is doing for whom, and at whose expense. A voice can be perfectly grammatical and still fail every one of the sixteen, if it is pitched at the speaker's ego. A voice can be technically halting and full of them, if it is *for* the listener.
This is also why the treatise puts *無羞聲* and *不怖聲* near the beginning of the list. Shamelessness (in the technical sense — not being beholden to offerings) and fearlessness (not being clenched by self-consciousness) are upstream conditions. If the speaker is secretly performing for an audience behind the audience, the other fourteen qualities will be subtly distorted. They will be deployed for *effect* rather than for *benefit*. And the ear — especially the ear of a child, a patient, a student, a follower, a partner — hears this distortion even when the mind does not name it. That is why the treatise is a treatise of *accomplishment* (成就), not of *technique*. Technique without inner freedom produces a voice that is technically impeccable and ontologically empty.
Why contemporary people need this
We live inside an audio economy that is, in a precise sense, the inverse of 字成就. The platform metrics reward *reach* (over 遍一切), *retention* (against 歡喜 — the listener should not tire, but the metric is engineered through variable-reward hooks, not through genuine worth), and *engagement* (a strange word that flattens 善友, 眾生根喜, and 應時 into a single number). The result is a culture in which almost everyone who speaks publicly is being subtly trained to optimize the *form* of the bodhisattva's sixteen while gutting their *substance*. Voices are widely heard, beautifully modulated, technically unimpeachable, and pointed at the speaker's brand.
This matters because the second-order effect is on listeners. When almost all the voices we hear are compromised in this way — *離不正聲* violated by euphemism, *不躁急聲* violated by speed-for-engagement, *無羞聲* violated by the constant subtext of monetization, *常流聲* violated by the algorithm's demand for novelty — the ear itself becomes unable to recognize what 成就 sounds like. People begin to mistake performance for care, and to experience authentic, audience-tuned speech as either naïve (when it is unpolished) or suspicious (when it is polished but uncorrelated with any apparent self-interest). The contemporary infodemic is not just a content problem; it is a *recognition* problem. We have lost the muscle for hearing when a voice is for us.
For this reason, the bodhisattva's sixty sounds are not a luxury for monastics or a relic of a pre-literate era. They are a diagnostic vocabulary. They let us notice, with some precision, where our own speaking has become hijacked by the metric system, and they give us a direction in which to train. They also let us recognize the rare voices we do encounter in which the qualities cluster — the teacher who is for-real-timed, the friend whose speech is free of the wincing subtext of self-presentation, the elder whose stories are unhurried. We can name what is happening in those voices, instead of just being vaguely moved.
And for those who teach — formally or informally, professionally or domestically — the sixteen are a way to audit one's speech without falling into the trap of turning speech into a self-improvement project (which would itself be a violation of *無羞聲*). The check is not "Am I sounding good?" but "Whose matter am I speaking? And is this sound landing where they are, or where I want to be?"
Common misreadings and clarifications
Misreading 1: "字成就 means mastery of language, vocabulary, eloquence." This is the most common flattening, and it is wrong in two directions. First, the treatise is explicit that it is collapsing *akṣara* into *śabda* — sound, not script; hearing, not reading. Second, the sixteen qualities are not about eloquence. Several of them (e.g., *不躁急聲*, *不增減聲*, *無羞聲*, *不怖聲*) are explicitly about the *absence* of flourish. The accomplished voice can be plain. The point is not how it sounds but how it functions for the listener.
Misreading 2: "The sixty sounds are a list of voice techniques to learn." This is the checklist error already named. The list is a phenomenology, not a curriculum. It is meant to make the reader sensitive to *what is going on* in speech, not to provide exercises. Treating *離不正聲* as "speak correctly" or *歡喜聲* as "sound warm" misses that the qualities are interdependent and rooted in the speaker's *intention and freedom*. The same external behavior can be empty or full depending on what is generating it.
Misreading 3: "This is only for bodhisattvas, so it doesn't apply to ordinary people." The bodhisattva figure in the *Mahāyāna-sūtrālaṃkāra* is a methodological ideal — the human being at the full ripening of the qualities being analyzed. The qualities themselves exist in seed in everyone who speaks. You can hear *應時聲* in a good bartender, *善友聲* in a long-married partner, *一切種成就聲* in a skilled storyteller at a dinner table, *不躁急聲* in a good therapist, *常流聲* in a grandparent reading to a grandchild. The treatise is naming what is already partially there and pointing toward its full ripening.
Misreading 4: "字成就 is about speaking more — being more articulate, more expressive, more articulate." It is often the opposite. *不增減聲* — neither adding nor subtracting — and *不躁急聲* — not rushing — are explicit constraints against more. The treatise is not endorsing prolixity. It is endorsing *fit*. Sometimes fit is silence. (Note: the treatise does not list silence among the sixteen sounds, but the qualities of *不增減* and *不躁急* imply it as the natural counterpart.)
Misreading 5: "The 'Tathāgata's matter' at the end is just a religious flourish — we can substitute 'truth' or 'the good' and keep the structure." This requires care. The treatise is making a strong claim that the perfected voice has a specific *telos*: the proclamation of the Tathāgata's true Dharma. This is not arbitrary. The point is that the voice's accomplishment is defined by what it serves, not by its internal polish. A voice that is technically accomplished but pointed at, say, political manipulation is, in this framework, the opposite of 字成就 — it is *正聲* (correct sound) corrupted into *不正聲*. The telos matters. One can secularize the telos (speak *for the benefit of the listener*, for *what is actually true*, for *what liberates*) without losing the analytical structure. One cannot drop the telos and keep the structure intact.
A final note on the choice of "word" versus "letter" in English translation. *Akṣara* — literally "imperishable" — sits at the root of the Greek *gramma* via Persian, and the *Mahāyāna-sūtrālaṃkāra*'s move to substitute sound for script is, in a quiet way, a refusal of the alphabetization of Dharma. It says: do not mistake the letter for the word, do not mistake the word for the sound, do not mistake the sound for the meaning. Each of these is a step of abstraction, and each step is a place where the living teaching can be lost. The treatise's sixty sounds are not the bottom of the chain but, in a sense, its return: the place where the abstraction meets an ear, and through that ear, a mind, and through that mind, the possibility of liberation. That is what is being "accomplished."