Dharma · AI Companion

Concepts × Canon Divisions

Aggregating 290k chunk-level core concept annotations across 123 texts (15.1M chars) by Taishō canon division yields a quantitative distribution of every concept across the canon — our first piece of computational Buddhology. Not a replacement for traditional doctrinal classification (panjiao), but a statistical mirror held up to it.

Finding: Emptiness exported a vocabulary; Yogācāra built a terminology wall

Traditional panjiao treats Prajñā (emptiness) and Yogācāra as the two wings of the Mahāyāna. Measured against the corpus, the two wings behave in completely asymmetric ways.

Emptiness-family concepts are the canon's lingua franca: 空 emptiness appears in 20 of 21 divisions (distribution entropy 0.80), 般若 prajñā in all 21 (0.80), 因缘 hetu-pratyaya in 20 (0.85). Ten core emptiness-family concepts average entropy 0.68 — the closest thing to a uniform distribution in the whole matrix.

Yogācāra terms barely leave home: 转依 āśraya-parāvṛtti occurs in just 2 divisions (entropy 0.05), 遍计所执性 in 5 (0.15), 藏识 ālaya in 10 (0.32). Eight core Yogācāra concepts average entropy 0.22 — the strongest “dialect zone” in the corpus.

The sharpest cut comes from the Chan records: emptiness-family concepts have 127 core links in the Recorded Sayings division; Yogācāra concepts have 11. That Chan speaks Madhyamaka rather than Yogācāra is now visible at chunk-level statistical resolution.

This ranks no school above another. Panjiao classifies by doctrinal register; the corpus stratifies by linguistic diffusion. That emptiness became the canon's common tongue fits its traditional self-description as the “shared dharma”; Yogācāra's self-containment fits its character as a precision scholastic system. What the machine view adds is that both characters can now be measured.

Density matrix: top 120 concepts × 19 divisions

Most universal (highest entropy)

因緣H 0.8520 div.
五陰H 0.8420 div.
方便H 0.8418 div.
三昧H 0.8420 div.
輪迴H 0.8318 div.
無明H 0.8319 div.
大乘H 0.8320 div.
第一義H 0.8217 div.
供養H 0.8218 div.
慈悲H 0.8121 div.

Terminology walls (lowest entropy)

轉依H 0.05瑜伽部
普賢行H 0.1華嚴部
遍計所執性H 0.15瑜伽部
無分別智H 0.16瑜伽部
依他起性H 0.22瑜伽部
大般涅槃H 0.25涅槃部
圓成實性H 0.26瑜伽部
畢竟空H 0.27釋經論部
解脫門H 0.29華嚴部
H 0.31阿含部
種子H 0.32瑜伽部
菩薩行H 0.32華嚴部

Division signatures (highest lift)

金剛三昧密教部×47.3
菩提道法華部×33.6
密教部×32.5
佛道法華部×26.5
法輪法華部×26.3
一闡提涅槃部×25.3
大般涅槃涅槃部×24.3
中觀部×22.5
三昧耶密教部×22.3
悉地密教部×21.6
加持密教部×21.1
中觀部×20.3
魔業般若部×19.3
般若波羅蜜般若部×18.8
一乘法華部×17.5

Method & limits

  • Data: per-chunk concept extraction over 46,545 chunks (no sampling, 96.7% coverage), extract×core tier only.
  • lift = per-chunk rate within a division ÷ canon-wide rate; entropy normalized over 21 divisions. Two tiny divisions (<60 chunks) excluded from the matrix.
  • Divisions follow Taishō/Xuzangjing philology; our corpus is a 123-text selection, not the full canon. Family samples are listed on-page; different samples shift numbers, not direction.
  • Statistics measure where language is used, not where doctrine is established — panjiao is a hermeneutic tradition, this page is corpus measurement; they mirror, not replace, each other.