一、這個概念在說什麼
When the *Mahāyāna-sūtrālaṃkāra* (《大乘莊嚴經論》) talks about a bodhisattva's "speech accomplishment," it is *not* talking about eloquence, charisma, public-speaking skill, or the capacity to dominate a room. It is talking about something far more specific and far more demanding: a kind of speech whose every layer — the speaker's inner conduct, the content taught, the words chosen, the very sounds emitted — is so integrated and purified that the speaking itself becomes a vehicle of awakening for those who hear it. That quality — speech that is *deśanā-sampad*, "說成就" — sits inside the bodhisattva's larger equipment (sambhāra) as one of the forces through which beings are ripened.
The *Mahāyāna-sūtrālaṃkāra* fixes the concept with an overarching frame of four markers: the speaker must be fearless before audiences (無畏), must be able to cut through the listener's doubts (斷疑), must be able to generate authentic faith in those who hear (令信), and must reveal the true principle of things (顯實). These four are the "skeleton" — every layer beneath has to be calibrated to produce them. Modern readers who already know the bodhisattva path should note what is missing: nothing here mentions persuasion, nothing mentions authority, nothing mentions metrics of engagement. The measurement is *whether the hearer is brought closer to truth*.
Underneath this overall frame, the treatise distinguishes three nested layers, each with its own list of qualities. They are not interchangeable, and the difference matters.
1. 說法成就 — Accomplishment in the teaching itself (the speaker's conduct + the dharma's content)
This first layer is about what the teaching does inside the speaker, plus what kind of content is delivered. It is expressed in nine conditions, which together tell us what kind of person can speak the dharma *at all*:
- 美語 (beautiful speech): When reviled or attacked, the speaker does not return harshness for harshness. The teaching begins with a refusal to mirror violence. - 離醉 (freedom from intoxication): When praised — or when possessed of family, beauty, wealth, status — the speaker does not become drunk on these things. Status does not inflate; flattery does not inebriate. - 無退 (no turning back): The teaching is not given out of partial application. The speaker does not slacken or quit halfway. - 無不盡 (no withholding): Out of stinginess toward the dharma, the speaker does not hold back. Nothing is withheld for leverage, for self-protection, or to keep an audience dependent. - 種種 (variety): The teachings are manifold and not repetitive — adapted to the wide range of beings encountered. - 相應 (correspondence): What is taught does not contradict either direct perception (現量) or reasoned inference (比量). In other words, the dharma content must cohere with what is verifiable. - 令解 (leading to comprehension): The formulation must actually unlock understanding; unintelligible speech is a failure here even if it is poetic. - 非求利 (not seeking gain): The speaker does not use material inducements to make the listener "believe." No payola, no bribes of belonging. - 遍教授 (universal instruction): The teaching is offered across all three vehicles — there is no sectarian hoarding of the dharma for one audience.
Already we can see that 說法成就 is mostly an ethics of the speaker. It is about the inner conditions under which speaking the dharma is permitted, possible, and trustworthy. Modern ears will hear it as something like "the speaker has done their own work first."
2. 語成就 — Accomplishment in verbal expression
If the first layer guarantees the *speaker and content*, the second layer addresses how the language is shaped so that it can do its work. Eight qualities:
- 不細 (not narrow): The speech reaches across the gathering, not reserved for a few. It is not coded as insider knowledge. - 調和 (harmonious, pleasing): It sounds good — refined, agreeable. - 善巧 (skillful): Each word and phrase is well-deployed. - 明了 (clear and luminous): Nothing is muddled. - 應機 (responsive to the listener): The speech matches the hearer's readiness, vocabulary, and concern. Different beings, different gear. - 離求 (free from seeking): No buried pitch, no covert bargaining. - 分量 (right measure): Neither overlong nor clipped — given in amount that can actually be received. - 無盡 (inexhaustible): The hearer is not left sated; there is a quality in the expression that opens rather than closes the appetite.
What this layer is really after is speech as a medium that can hold transformation. Modern ears will recognize in this list almost every quality one would want in a careful translation, an essay that changes how you see things, or a teacher who can actually be heard.
3. 字成就 — Accomplishment in the voice itself
The third layer goes one step further down — from content, to language, to the actual sonic quality of the voice and its function in space and time. Sixteen kinds of "voice" are enumerated. They fall into recognizable clusters:
Purity of the voice itself: *離不正聲* (free from wrong sound — no affectation, no distortion); *不毀呰聲* (not disparaging, not crude); *不增減聲* (neither exaggerated nor understated); *不躁急聲* (not rash or rushed); *無羞聲* (not shame-laden); *不怖聲* (not fearful).
Right relationship to time and audience: *應時聲* (timely — speaks when the moment is right); *歡喜聲* (gladdening); *隨捨聲* (giving freely); *善友聲* (like a good friend's voice); *常流聲* (continuously flowing, not stopping); *遍一切聲* (reaching all — both near and far disciples); *嚴飾聲* (adorned, dignified).
The two highest capacities — these are the ones to notice carefully: *滿足聲* (fulfilling voice): one sound contains innumerable sounds, the speaker can say one word and a thousand meanings are heard by those ready for them; *眾生根喜聲* (voice pleasing to beings' faculties): one phrase reveals innumerable meanings. Then closing the list, *一切種成就聲* — the voice that accomplishes every kind of benefit.
These last two are the powers at which the layer is aiming: the speech is not merely informational — it is *generative*. A single utterance can meet a thousand different listeners where they actually are. This is what the treatise means when it says 說成就 sits at a "most supreme" position among the forces by which beings are ripened — it is not because the speaker is talented, but because the speech has been refined enough to act like a *function*, not a performance.
Why these three layers must not be flattened
Contemporary readers will be tempted to collapse all three into "good communication." The treatise itself warns against this. The point is not that 說法成就 is "deeper" than 語成就 which is "deeper" than 字成就. It is that they cover three different surfaces of a problem: whether the speaker is fit to teach (ethics); whether the teaching is fit for the listener (pedagogy); whether the voice is fit to carry the teaching across time and space (vocal/somatic/spatial). A teacher who is ethically clean but mumbles fails at the third layer; a teacher with a beautiful voice but who teaches from ego fails at the first; a teacher with great pedagogy who is only narrow and elite fails at the second. The three must move together.
Where this sits in the bodhisattva path
說成就 is not an isolated virtue. It belongs to the bodhisattva's equipment-for-ripening (成熟眾生), and the *Mahāyāna-sūtrālaṃkāra* explicitly pairs it with 成熟眾生力 — the *power to ripen beings*. It is what the bodhisattva gets praised by the buddhas for; it is the language-side of compassion-in-action. And the fruits named in the treatise are not rhetorical: 善說 can lead beings to come-toward (歸向), enter (趣入), be tamed (調伏), accomplished (成就), abide (安住), awaken (覺悟), and finally be liberated (解脫). It can gather merit, lead to higher rebirth, produce *vyākaraṇa* (受記, the Buddha's prediction), and culminate in the full awakening of a tathāgata's knowledge (成就如來智). The arc runs from a single utterance to full Buddhahood. That is how seriously this text takes the quality of how a bodhisattva speaks.
二、生活裡的走查
Now let us walk through 說成就 with someone we know well — say, a mid-career professional in a large company who has been asked to mentor a junior colleague, or, in another register, a parent trying to explain something hard to a teenager, or, in a third, a content creator with an audience they did not ask for. The classical categories will hold up better than you might expect.
Start at 說法成就 — the question is not "what should I say" but "am I the right person to be saying it." 美語, before anything else, is the colleague-mentor who, when the junior pushes back rudely or the teenager yells, does not return sharpness. Not because they are suppressing; because the speech-act they are doing — if it is to be 說成就 at all — requires that they not turn the conversation into a contest of wounds. 離醉 is the mentor who has just been promoted and notices that the junior, hungry for approval, is now treating every word as gold. The mentor's job is to notice that the new gold-leaf on their authority has nothing to do with what they are trying to say. 無退 is the parent who has explained the same thing five times and feels their patience dying; the bodhisattva, and the treatise, do not say "quit when tired"; they say "do not turn back." 無不盡 is harder and easier to see: it is the colleague who, when asked a real question, holds back the most useful part because "they should figure it out themselves." That is 慳法 — stinginess with the dharma — and it disqualifies the speech even if everything else is in order. 種種 is the recognition that not every learner needs the same analogy. 相應 is harder to fake: a mentor who tells a junior to "just trust the process" without that process being grounded in what can be perceived and inferred is failing 相應 — teaching what cannot meet reality. 令解 is the grim humility that beautiful phrasing which the listener cannot *comprehend* is no teaching at all. 非求利 is a direct hit on a recognizable contemporary failure mode: any speech which is bent toward followers, money, status, prestige, a book deal, an audience for some other project — that speech becomes 說法成就 no longer, no matter how luminous. 遍教授 is the refusal to teach only the students who will flatter the teacher back.
What is striking when you actually try to map this onto a real mentoring conversation is that the first layer collapses almost entirely into the teacher. It is not a set of "speech tips." It is a set of *permissions, conditions, and disqualifications on the speaker*. The classical text is correctly refusing the modern assumption that a clear message with a clear voice is enough — it begins with the inner state of the one talking.
Move to 語成就 — now the speaker is permitted to speak; what language do they use? The first four (不細, 調和, 善巧, 明了) are the basic craft of saying something clearly and well. 不細 alone can transform a meeting: a status update addressed to the whole team in insider acronyms that only three people understand is failing 語成就 — the speech is *narrow*. 調和 is not "people-pleasing"; it is the quality of utterance that listeners do not feel rubbed against. 應機 is the recognition that the same content needs to be repackaged for a nervous teenager, an exhausted colleague, a defensive spouse. 離求 is the hidden one: an instruction that subtly positions the speaker as indispensable, that presents itself as generous but is actually a recruitment — fail. 分量 is the modern plague's remedy: meetings that could be five minutes becoming fifty. 無盡 — and here the treatise is doing something subtle — is not "the speech is bottomless." It is that the listener, at the end, is not satiated. They want to know more. They feel *opened*, not filled. Most pedagogy fails at 分量 and 無盡 at the same time: it overstuffs until nothing is alive.
Move down to 字成就 — and now the question is the texture of the voice itself, in the room. This is the layer where physical presence, breath, register, timing, all of it, enters. 應時聲 is the parent who *waits* until the teenager is no longer raging before they speak; the mentor who does not deliver the crucial feedback ten minutes before the junior's presentation. 不躁急聲 is the mentor who notices they are speeding up because they themselves are anxious, and slows down. 不增減聲 is the one that contemporary speech increasingly fails: we have learned to either exaggerate everything (the great-crisis tone of news media, social media theater) or underplay everything (the cynical shrug of irony). 嚴飾聲 — dignified adornment — is what keeps a piece of teaching from being merely correct and unremarkable. 無羞聲 and 不怖聲 tell us what kind of fear the bodhisattva's voice is free from: the fear of being laughed at, of being seen as foolish, of saying something the audience might reject. 善友聲 is the surprise of the list: the bodhisattva's voice *sounds like a good friend*. This is not sentimental. It is a technical specification: the voice should carry the specific warmth of someone who is on your side.
The two transformational capacities, 滿足聲 and 眾生根喜聲, ask: can one utterance land differently in different listeners? A practiced teacher says one sentence and a beginner hears "be kind to yourself"; a more advanced listener hears "do not mistake the impermanent for the self"; an advanced practitioner hears a teaching on emptiness. *One sound, many meanings.* Contemporary communication theory has rediscovered this as audience-adaptive speech, but the treatise predates and exceeds the contemporary version by saying it must be unforced — part of the voice's own accomplishment, not a performance trick. The 16th voice, *一切種成就聲*, is the final umbrella — voice that accomplishes every kind of benefit one might hope speech could accomplish. The treatise is not modest.
Where the modern mapping goes wrong most often. Three traps. The first is collapsing the three layers into one — taking the list as if it were simply "good speaking tips" and missing that 說法成就 is *ethical*, 語成就 is *pedagogical*, and 字成就 is *vocal-somatic*. The second is reducing 說成就 to performance metrics: views, ratings, watch time. The treatise refuses this completely; 非求利 and 離求 are non-negotiable. The third is the trap the modern listener is most likely to fall into unconsciously: assuming the bodhisattva's speech should be *calm, neutral, without personality.* 歡喜聲, 嚴飾聲, and 善友聲 together say the opposite — the voice is permitted to be glad, adorned, warm. What it must not be is performative or unattuned. The discipline is not flatness; it is alignment between the inner state and the outer voice.
A worked example. Imagine a senior engineer explaining to a junior why a particular architecture must be redone (not what the junior wants to hear). The internal state matters first — is the senior speaking out of resentment that they have to redo something? Out of pleasure in correcting? Out of genuine attempt to develop the junior? 離醉 and 非求利 are already at work. Now the language: is the explanation either narrow (in-jokes, insider acronyms the junior hasn't earned yet — failing 不細), or appropriate to the junior's level (應機)? Now the voice: is it bored (failing 歡喜聲), rushed (不躁急聲), exaggerated (不增減聲)? If all three layers hold, the junior leaves *disoriented but with something to think about* (無盡) and able to *actually understand why* (令解). If any of the three layers fails, the meeting is experienced as either a humiliation, a non-event, or a guilty dramatization. Modern technical communication is mostly failing all three at once, and the *Mahāyāna-sūtrālaṃkāra* is exactly the sort of classical text that lets us name *why*.
三、當代人為什麼需要它
We live inside an information economy in which speech has been dissociated from almost everything the *Mahāyāna-sūtrālaṃkāra* insists it must be connected to. Speech is now optimized for what it can extract — attention, affiliation, revenue, micro-status, persuasion that someone has not consented to. The classical list reads almost as a diagnostic of the contemporary failure: 美語 is the norm-violation of a public sphere in which retaliation is normalized; 離醉 is what almost no public figure of any size can demonstrate; 無不退 and 無不盡 are blocked by the structural incentives of partial disclosure (you build audience by *teasing* what you will explain later); 種種 is broken by personalization algorithms whose economy is sameness; 應機 is broken by mass address that treats the audience as a demographic rather than as beings of differing capacities; 非求利 is structurally rare; 遍教授 is impossible in a recommendation-engineered information diet. Almost every contemporary condition of speech fails 說成就 on multiple axes at once.
This is not a complaint about the modern world. It is the reason the classical category is useful: it gives us a vocabulary for a problem we already feel but cannot easily name. When we say a teacher has "presence," we are reaching toward 字成就. When we say a colleague's explanation "actually landed," we are pointing at 令解 and 應機. When we say a podcast host "doesn't seem to be selling anything" and that is why we trust them, we are invoking 離求 and 非求利. The concept provides precision to distinctions we are already making instinctively.
It is also useful at the level of inner life. Most contemporary people who speak publicly about anything important — to a team, a class, a parent group, a partner — silently worry: *am I performing? Am I selling myself? Do I actually know what I'm saying? Am I reaching anyone?* 說成就 turns these private worries into a structured set of practices: cultivate the speaker so the speech is permitted, refine the language so the speech can carry, refine the voice so the speech arrives and lingers. In a culture that has outsourced both training and reflection about speech to metrics and templates, this three-layer interiority is genuinely valuable. It is also worth noticing that the three layers correspond almost exactly to three things modern attention research has confirmed matter: the speaker's credibility (說法成就), the structure and clarity of the message (語成就), and the affective and prosodic delivery (字成就). This alignment with contemporary work on, for instance, the trustworthiness cues that listeners pick up prosodically is suggestive — but we should be careful. The classical text does not claim to be optimized for engagement metrics, and there is genuine research indicating that some elements (e.g., the warmth of 善友聲, the dignity of 嚴飾聲) sit in a region where contemporary behavioral science has not yet given us a settled mechanistic account. The convergence is suggestive, not proof, and it would distort the concept to read it as "what science has now confirmed." It is, rather, what two very different epistemic traditions have independently converged on: speech is layered, and the layers are not interchangeable.
The deepest reason a contemporary person might need this concept is more personal. We are also, mostly, listeners. We are drowning in speech that is either formulaic or persuasive but not aimed at our actual freedom, and we know it. 說成就 names what speech aimed at the listener's liberation actually requires, and is therefore a way to recognize what we are missing in the speech around us. It is also a way to recognize, and honor, the rare speech we do encounter that is genuinely doing this work — and to ask, of our own speech, what would it take to be more of that and less of the substitutes.
四、常見誤讀與澄清
1. "說成就 = eloquence or rhetoric." The most common flattening. The treatise explicitly refuses this: 說法成就 is bound to ethics (離醉, 非求利), 語成就 is bound to pedagogy (令解, 應機), and 字成就 is bound to function (滿足聲, 眾生根喜聲). Reduce it to eloquence and you lose the entire ethical and pedagogical spine. A person can be a magnificent orator who is failing 說成就 on every internal axis.
2. "The three layers are just progressive depth." They are different surfaces of the same problem, not levels of "advanced." A beginner teacher must work on the first layer (ethics + content); a polished teacher must still work on the same first layer, because the ethics can be lost at any level of fame. They are not sequential stages of progress; they are simultaneous requirements.
3. "美語 means sweet or flowery speech." 美語 in 說法成就 is *not* a speech-style. It is a *non-retaliation*: when insulted, the teaching continues without matching cruelty. Treating it as "eloquent" loses the entire meaning. 美語 is an ethical stance under pressure, not a linguistic achievement.
4. "離醉 = humility." Not quite. It is specifically the freedom from intoxication by one's own advantages — looks, lineage, wealth, accomplishments. Humility is one expression of it; indifference to flattery is another. The point is not low self-regard; it is the absence of any sense of being *higher* than the listener because of contingent attributes.
5. "應機 = being accommodating." No. 應機 means matching the speech to the listener's actual *capacity*, including their readiness to hear something difficult. A accommodating speech that tells people what they want to hear is failing 應機 if the listener's actual need is to be unsettled. It is the opposite of a people-pleasing flexibility; it is a calibrated firmness.
6. "字成就 is about beautiful voice / voice training." Only at the surface. The sixteen sounds together specify a voice that is *functionally aligned*: it is timely, free, joyful, dignified, friendly, not-rushed, and — at the highest two — capable of one-utterance-innumerable-meanings. It is closer to vocal integrity than to vocal beauty. A singer with a mellifluous voice who cannot reach a beginner and an advanced practitioner with one sentence is failing 字成就 at its highest entries.
7. "無盡 means the speech should be long." The opposite. 無盡 is the quality by which the listener is not *sated*. It is sometimes achieved in two sentences. Long, exhaustive lectures are usually failing 無盡 because they fill the listener rather than open them.
8. "The bodhisattva's speech is detached, neutral, selfless." The list names 歡喜聲, 嚴飾聲, 善友聲. The voice is permitted gladness, dignity, warmth. It is restrained from performance, attachment, manipulation, fear — but it is not flat. The contemporary mistake is to read detachment as the absence of affect; the classical list is more specific. It is the absence of *self-concern in* the affect. Joy is allowed. Adornment is allowed. Warmth is allowed. What is excluded is *them being for the speaker's sake*.
9. "說成就 is about converting people." Read carefully: the four guiding principles (無畏, 斷疑, 令信, 顯實) include 令信 — generating faith — but it is bracketed by 斷疑 (cutting the *honest* doubts) and 顯實 (showing the real). It is not the engineering of belief; it is the removal of doubt so that what remains is faith in what can be verified. Conversion-rhetoric, which seeks to manufacture faith in what cannot be verified, fails at 相應 and 顯實 simultaneously.
10. "We can't operate at this level — it's a high-stage bodhisattva thing." The treatise does not say this. 說成就 is positioned as part of the equipment that ripens beings — it is something that bodhisattvas at various stages cultivate. Each of the qualities is locally *practicable* by an ordinary person who wants to take their speech seriously: do not retaliate when insulted; do not get drunk on praise; do not turn back mid-teaching; do not withhold; vary your teaching; do not contradict evidence; make yourself understood; do not sell; teach across differences. These are available, now, to anyone willing to do the work the layers describe. The concept is *aspirational in its full form, practicable in its parts*, and that combination is exactly what makes the category useful to a contemporary reader who does not identify as a bodhisattva but who does, perhaps, want to take their speech seriously as a craft and as an ethical practice.