Concepts × Canon Divisions
Aggregating 290k chunk-level core concept annotations across 123 texts (15.1M chars) by Taishō canon division yields a quantitative distribution of every concept across the canon — our first piece of computational Buddhology. Not a replacement for traditional doctrinal classification (panjiao), but a statistical mirror held up to it.
Finding: Emptiness exported a vocabulary; Yogācāra built a terminology wall
Traditional panjiao treats Prajñā (emptiness) and Yogācāra as the two wings of the Mahāyāna. Measured against the corpus, the two wings behave in completely asymmetric ways.
Emptiness-family concepts are the canon's lingua franca: 空 emptiness appears in 20 of 21 divisions (distribution entropy 0.80), 般若 prajñā in all 21 (0.80), 因缘 hetu-pratyaya in 20 (0.85). Ten core emptiness-family concepts average entropy 0.68 — the closest thing to a uniform distribution in the whole matrix.
Yogācāra terms barely leave home: 转依 āśraya-parāvṛtti occurs in just 2 divisions (entropy 0.05), 遍计所执性 in 5 (0.15), 藏识 ālaya in 10 (0.32). Eight core Yogācāra concepts average entropy 0.22 — the strongest “dialect zone” in the corpus.
The sharpest cut comes from the Chan records: emptiness-family concepts have 127 core links in the Recorded Sayings division; Yogācāra concepts have 11. That Chan speaks Madhyamaka rather than Yogācāra is now visible at chunk-level statistical resolution.
This ranks no school above another. Panjiao classifies by doctrinal register; the corpus stratifies by linguistic diffusion. That emptiness became the canon's common tongue fits its traditional self-description as the “shared dharma”; Yogācāra's self-containment fits its character as a precision scholastic system. What the machine view adds is that both characters can now be measured.
Density matrix: top 120 concepts × 19 divisions
Most universal (highest entropy)
| 因緣 | H 0.85 | 20 div. |
| 五陰 | H 0.84 | 20 div. |
| 方便 | H 0.84 | 18 div. |
| 三昧 | H 0.84 | 20 div. |
| 輪迴 | H 0.83 | 18 div. |
| 無明 | H 0.83 | 19 div. |
| 大乘 | H 0.83 | 20 div. |
| 第一義 | H 0.82 | 17 div. |
| 供養 | H 0.82 | 18 div. |
| 慈悲 | H 0.81 | 21 div. |
Method & limits
- Data: per-chunk concept extraction over 46,545 chunks (no sampling, 96.7% coverage), extract×core tier only.
- lift = per-chunk rate within a division ÷ canon-wide rate; entropy normalized over 21 divisions. Two tiny divisions (<60 chunks) excluded from the matrix.
- Divisions follow Taishō/Xuzangjing philology; our corpus is a 123-text selection, not the full canon. Family samples are listed on-page; different samples shift numbers, not direction.
- Statistics measure where language is used, not where doctrine is established — panjiao is a hermeneutic tradition, this page is corpus measurement; they mirror, not replace, each other.