🛡️ TailGuard corpus report

SNAPSHOT  training-corpus-v1 · TailGuard v0.1.0

MetricCurrent
Documents1500
Tokens325426
Vocabulary27072
Vocab / 10k tokens831.9
Type-token ratio0.0832
Bigram diversity0.5641
Trigram diversity0.8908
Mean burstiness0.457
Near-duplicate rate0.0
Mean synthetic score16.2/100
Flagged synthetic0.1%

Synthetic-score distribution

0–9
72
10–19
1245
20–29
166
30–39
12
40–49
3
50–59
1
60–69
1
70–79
0
80–89
0
90–99
0

Most-repeated 5-grams (boilerplate check)

Distribution-level corpus health, per "Contamination Detection Is Not Collapse Monitoring" — per-document detection misses collapse; these metrics catch it. habsburg-ai.vercel.app