Skip to content

About

The load-bearing vocabulary of Claude: cluster analysis of GitHub pull requests

Resources

Stars

107 stars

Watchers

2 watching

Forks

Repository files navigation

The load-bearing vocabulary of Claude

GitHub pull request descriptions, grouped by the words they are written with rather than by anything they were told to look for: ten ways of writing, and every description belongs to one of them. One of the ten was 0.7% of the corpus at the start of 2025 and is 39% of it by the middle of 2026.

louisabraham.github.io/load-bearing

file what it is
fetch_day.py ten requests a day to GitHub's search API, one data/days/YYYY-MM-DD.jsonl. Standard library only.
analyze.py reads the days, groups them into whole weeks, fits the model, writes analysis.js and model.js. Needs numpy, scipy and numba.
index.html reads analysis.js. One board, one screen: the figures, the stack, a word's own history, the thousand words. No build step. Open it.
detector.html reads model.js. Paste a description, or a GitHub link, and the same fit says whether it is in the arriving cluster. Runs in the page; §7.
model.js the whole of the fit written down — every word, all ten numbers — in 304 kB. Nothing but the classifier reads it.
tests/ what the two pages must keep doing, driven in a real browser. Every test is a bug one of them once had.
.github/workflows/daily.yml does all of the above daily, commits the corpus here, publishes the page to gh-pages.
pip install numpy scipy numba
export GITHUB_TOKEN=$(gh auth token)

python fetch_day.py                  # yesterday, ten requests
python fetch_day.py --backfill 30    # and the last 30 days, if missing
python analyze.py                    # ~50 s on twelve cores
python analyze.py --selftest         # the invariants, on synthetic data
open index.html                      # the board
open detector.html                     # the same fit, asked about one text

pre-commit install                   # optional: ruff and the html formatter, on commit
uv pip install pytest-playwright && pytest tests -q

Current state: 603 collected days, 595 of them in 85 whole weeks (2025-01-06 to 2026-08-17), 461,121 descriptions, 51,079,244 word appearances, 19,798 words above the floor.


1. Why GH Archive cannot be used

The natural source is the public archive of GitHub's event stream, and it stopped working. Since mid-2025 the feed carries almost only PushEvent — a complete hour of 2024-08-12 holds 13,555 IssueCommentEvent against 86 for the same hour of 2026-08-10 — and pushes carry no text, GitHub having removed the commit array in October 2025. The cause is upstream in the Events API: #310 has been open since July 2025 with no maintainer reply, and the same gaps appear in OSSInsight, which reads the API directly. No mirror repairs it, because they all read the same feed.

This was found the hard way. An earlier version of this project, built on the archive, reported load-bearing in 17 documents. That was wrong by a factor of 158: the comments had disappeared from the feed, not from GitHub.

2. How the data is collected

What works is GitHub's search API, for one reason: created: accepts timestamps and not only dates, so a window can be minutes wide and every response carries the full body text.

Ten five-minute windows a day, one drawn from each 2.4 hours of it. The ten starts are drawn to the second and one per block, which is not fussiness: they used to be multiples of five minutes, which is exactly the granularity a cron schedule fires on, so every window opened at an instant when scheduled automation opens pull requests. Blocks also keep the ten from clumping and make it impossible for two to overlap. The draw is seeded on the date, so the whole corpus is reproducible from its dates alone, and each day is one immutable file of about 1.4 MB, committed and never rewritten. The repository's history is the history of the sample.

Two filters go into the query itself and together take a page from 43 usable descriptions to 97: four Apps excluded by name — pull, dependabot, renovate, github-actions, which are 90% of App-authored bodies — and empty bodies excluded, 45% of all pull requests. There is no emptiness qualifier in the search API; requiring any one of ten function words in the body does it exactly.

One honest limit. A day of 2026 holds some 460,000 pull requests matching the query, a five-minute window about 1,250, and a page is a hundred — so a window is truncated to its earliest hundred and this samples rather than enumerates. Nothing here could enumerate a day: the search API returns at most 1,000 results per query however many matched. Uniform placement means this is not a bias in time; the effective width just narrows as GitHub gets busier. A day can also come in short, and 116 of the 603 do — mostly early 2025, when a window did not always fill its page, but two days hold 900 because one of their ten windows returns nothing at all and returns nothing again when asked twice. Those are holes in GitHub's own index, and the day is written short rather than patched.

The corpus and the site live on different branches. A published Pages site may be no larger than 1 GB and the corpus grows 1.4 MB a day, so the daily run commits the day here and pushes only index.html, analysis.js and .nojekyll — a quarter of a megabyte — to gh-pages. The corpus keeps its history because its history is the point; the site does not need one.

3. How the data is cleaned

A word is a run of letters, digits, slashes, hyphens and underscores containing at least one letter, so load-bearing, snake_case, --all-targets and src/main survive whole. No stemming, no n-grams, no stopword list. Links collapse to their domain and HTML tags are taken whole, because splitting on punctuation first put bugbot](https and href among a component's most characteristic words. The em dash is the one deliberate exception to requiring a letter, and it earns it: 0.2 appearances per 10,000 words in early 2024 against 123 in mid-2026. Median description: 65 words.

What gets thrown away. Accounts that are not people, by the shape of the login — anything ending [bot] or -bot, plus copilot — which is 3,784 accounts and 13.2% of collected rows. Identical word sets within a week, because one ordinary human account posted 147 copies of one sentence in a fortnight. And no author may contribute more than three descriptions to a week, which catches mass-produced text from accounts that look human and applies to humans on the same terms, which is why it is a cap and not an exclusion.

One floor on a word, and it counts people

A word is in the vocabulary when 50 distinct accounts have written it. That is the only floor. There were three — 45 appearances, 25 descriptions, 20 accounts — and two of them were doing nothing this one does not do better, because counting appearances cannot tell a shared word from one document written two hundred times:

word appearances descriptions accounts
store-path 242 242 2 dropped
mq 569 533 36 dropped
load-bearing 1,011 905 848 kept
seam 1,849 1,247 1,135 kept

A word 848 people reached for is a word; a word in 242 descriptions from 2 accounts is one document written 242 times. The number is set on a property of the method, not on the answer: it is the least restrictive floor at which two independent fits agree on half of their top twenty words. Agreement rises with the floor all the way up, so there is no optimum to find — only a rate of return, and a rule that picks a point on it for a stated reason. It costs coverage: 19,798 words of the 2.1 million in the corpus, where the old three floors kept 26,113.

Whole weeks only. Seven days of ten windows is 7,000 descriptions collected and about 5,300 after the filters, so weeks are the same size by construction and need no cap. Part-weeks at either end are dropped outright, which matters daily — collection runs each morning, so the newest week is almost always half-collected, and it is the week everything leans on.

4. What the model is

Each of k ways of writing is a fixed distribution over the vocabulary, and every description is assigned to exactly one of them: the one it is closest to, under the divergence that belongs to word counts.

$$z_d \;=\; \arg\min_c \; n_d \, \mathrm{KL}(p_d \,\|\, W_c), \qquad W_c \;\propto \sum_{d\,:\,z_d = c} x_d$$

Each centre is the middle of what it was given — that cluster's KL-centroid. This is k-means with KL in place of squared distance, and the $n_d$ weight is the only trace of counting left in it: a long description pulls its centre harder than a short one. Nothing is ever evaluated as a divergence, because $x_d \cdot \log W_c = -n_d(\mathrm{KL}(p_d | W_c) + H(p_d))$ and $H(p_d)$ does not vary with $c$, so the nearest centre is the largest $x_d \cdot \log W_c$ and the assignment step is one sparse product against the corpus.

There is no t anywhere in that. One set of centres covers the whole window, so the fit has no per-week parameter — nothing that could describe a trend and no freedom to place one. Every curve the page draws is attribution instead: each description placed by its words alone, the weeks counted up afterwards. If a way of writing rises, the rise is in what people wrote, because there is nowhere else for it to be.

5. How the model is trained

Greedy k-means++ under KL, then Lloyd's algorithm to an exact fixed point — stop when no description changes hands, so there is no tolerance to choose and no pass count to guess. Eight fits from eight seeds, and the cheapest is published. The restarts are not there to find a better answer: cost correlates +0.03 with the share the page reports. They are there so the daily job publishes something.

What the page claims is that the component arrived, so two thresholds check it rather than select it: under 2% of the first eight weeks, at or above 20% of the last eight. Picking the biggest component says nothing about whether it arrived. If a batch fails the check, it runs again from fresh seeds; if four batches fail, nothing is published and the job stops. That retry would condition the fit on its own check, which is why the evidence is the rate at which unconditioned fits arrive: 31 of 32 single fits of this corpus, and in 1 of the 32 the leading component came out mixed with another. Where exactly it ends is one fit's answer, and SEED is listed below for that reason.

The selftest runs before every publish and stops the job if it fails: the centres are distributions, the weekly counts are whole numbers that reconstruct each week's total, and a planted way of writing is recovered from synthetic data — 0.000 to 0.350 at the week it was planted, though the model has no way to represent time.

6. How the results are displayed

The component shown is the largest across the last four weeks — a month rather than a week, so the subject of the page does not turn on which of two close components led across one of them. "Still growing" is read off the data rather than typed into the markup: a least-squares line over the last 12 weeks, currently +1.2 points a week.

The words are ranked by counting, not by the fit. The assignment is hard, so every appearance belongs to exactly one component and the counts partition it:

$$\mathrm{ratio}(v) = \frac{x^{\,\text{in}}_v \big/ N^{\,\text{in}}} {\left(x^{\,\text{out}}_v + \tfrac{1}{2}\,\texttt{MIN\_AUTHORS}\right) \big/ N^{\,\text{out}}}$$

Each side is divided by its own size, which makes this a ratio of two frequencies rather than of two counts. The pseudo-count in the denominator is the difference between a ranking and a lottery. Without it the top of the list was decided by counts of two to seven in 42 million: a word written three times outside beat one written 158 times, on a difference no larger than its own noise, and both are too rare for anyone to have noticed. Half of MIN_AUTHORS is the fewest appearances a word in this vocabulary can have, halved — the honest prior for written outside less often than can be measured. A word never written outside then scores in proportion to what it was written inside, so the top is ordered by frequency among a component's exclusive words rather than by the accident of a tiny divisor.

The comparison is against everything that is not this component, not against the whole corpus, which would compare the component against itself. load-bearing is written 929 times inside and 82 outside — 39×, the top of the list — and size on the page follows the logarithm of the ratio, from there down to 5× at the thousandth word.

Choosing a word replaces the chart with that word's own history as a rate: appearances per million words written that week. The corpus is not the same size from one week to the next — 380,404 words in the thinnest, 1,406,687 in the fattest — so a curve of raw counts would draw the corpus growing wherever it draws the word arriving.

7. The model is also a classifier

Nothing is added to the fit to make one. The assignment step is $\arg\max_c x_d \cdot \log W_c$, and $x_d \cdot \log W_c$ is the log-likelihood of the description under a multinomial that draws every word from $W_c$. Normalise the ten and they are a posterior:

$$P(c \mid x) \;=\; \frac{\prod_v W_c[v]^{\,x_v}}{\sum_{c'} \prod_v W_{c'}[v]^{\,x_v}}$$

which is multinomial naive Bayes with these centres as its class-conditional distributions. SMOOTH is what makes it usable on text the fit never saw: no centre gives any word probability zero, so one unexpected word cannot zero a whole component. There is no $\pi_c$ in it — weighting by how big each component is would answer what does a description of 2025 and 2026 usually look like over the top of the question that was asked, which is about the text in the box. detector.html does that arithmetic in the browser and asks one thing of it: is this the component that arrived, or one of the other nine? It reads vocabulary, so it can say a text is written like that cluster and never who wrote it.

How model.js fits in 304 kB

The classifier needs the whole of $W$ — ten numbers for each of 20,309 words — which is 4 MB of JSON. It is written as text instead, one character at a time, and the page rebuilds it in one pass.

The vocabulary is a trie, flattened: 174 kB of words in 78 kB. Sorted, each word stores how much of its predecessor it repeats and then the rest of itself, which is the tree written out in the order a walk of it visits — the shared prefix is the path already climbed and the suffix is the branch. The repeat count is a capital letter, and that is why the words need no separator between them: a word is lowercased before it is ever counted, so the capital that opens one is also what ends the one before it.

Each weight is one character from an alphabet of 92. What is stored is not $\log W$ but

$$E \;=\; \log\!\left(1 + M/\texttt{SMOOTH}\right), \qquad \log W_c[v] \;=\; \text{floor}_c + E[c,v]$$

where $M$ is the appearances a component holds of the word. $E$ is exactly zero wherever the word is absent — a quarter of the entries — and $\text{floor}_c$ is one number per component that the page adds once per word of the text. So only $E$ has to be written down, and it runs from $\log(1 + 1/0.01) = 4.615$ for a word written once to 17.978 for the commonest word in the corpus.

Code 0 means the component never wrote the word. Codes 1–80 stand for a value on their own: a grid over $[4.615, 10]$, one step every 0.068 nats. The 11 codes left over stand for no value at all. Each of them says this number is in the top of the range and it needs the next character too, and the pair then picks one of $11 \times 92 = 1{,}012$ levels over $(10, 17.978]$ — one step every 0.0079 nats, nearly nine times finer than the single-character grid.

Eleven, because eleven is what is left: 92 codes, one for absent, eighty for the grid that fits in one character. The two halves trade against each other — every code given to the single-character grid is one fewer block of 92 in the two-character one — and eighty, with the split at 10 nats, is the pair that came out smallest for the accuracy it keeps. The top of the range is worth the second character because that is where the commonest words are, and an error there is paid once per appearance.

load-bearing is the whole scheme in one word. Its ten counts are 945, 166, 5, 0, 5, 0, 3, 1, 1 and 0, so it needs all three kinds of code:

component appearances $E$ code written as
0 945 11.456 83, then 1 — a pair v#
2 5 6.217 24 :
3 0 — 0 !

Code 24 stands for $E = 6.183$ against a true 6.217, and that gap of 0.034 nats is the largest error anywhere in the file. The word's ten numbers are the eleven characters v#o:!:!3##! — ten codes, one of which needed a second character.

What the compression costs, against the exact centres over the whole corpus: no entry is more than 0.034 nats out, and 0.054% of descriptions land on a different component — all of them ties, with a median gap of 0.04 nats between their top two against 15.6 across the corpus. analyze.py --selftest reads the file back and fails the build if either half stops reconstructing the fit. The page also has its own copy of the tokeniser, in JavaScript, and nothing structural keeps the two in step: tests/test_detector.py runs both over the same strings and fails if they ever disagree.

8. The arbitrary choices

Everything above is either measured or a judgement call. These are the numbers that could have been different, and one of them was chosen by looking at the answer.

constant value how it was chosen
K 10 chosen on the outcome — see below
SEED 0 consequential — the seed moves the headline; §5
WINDOW_S 300 s consequential — see below
MIN_AUTHORS 50 measured — the least restrictive floor at which two fits agree on half their top twenty; §3
MAX_PER_AUTHOR 3 arbitrary
N_INIT 8 insurance — a single fit publishes 31 times in 32, so this is margin
LEAD_WINDOW 4 weeks judgement — "a month", to stop one week deciding the subject
LEAD_START, LEAD_END 2%, 20% round numbers, wide margins, and they only check
WORDS_LEAD 1000 arbitrary round number
BOT_SUFFIX, BOT_LOGIN [bot], -bot, copilot judgement — what a login says is not a person
ESCAPE_AT, SPLIT 80, 10 nats measured — the pair that costs 0.054% of assignments at 304 kB; §7

K = 10 was chosen on the outcome, but from a window rather than a preference. Below ten the component is a mixture: at k = 8 the leading component's own top twenty carries WebKit, nixos and CSS vocabulary in 7 of 32 fits, and the arrival check fails in 10 of them. Above fourteen the component splits until the pieces fall under the 20% the check asks for — 6 of 8 fits arrive at k = 16, 1 of 8 at k = 24. Ten and twelve are both inside the window; ten keeps the title word at the top of the list and reports a fuller share, twelve makes the arrival unanimous. Cost cannot settle it: training cost falls with every added centre and held-out cost is still falling at k = 64, so the corpus would happily support far more clusters than a reader can look at.

WINDOW_S = 300 is not the sample size it looks like. Five minutes was chosen so a window would fit in one page of a hundred results. It does not, in 2025 or 2026, so the sampler truncates rather than enumerates and the width is doing almost nothing — it is a floor that guarantees a full page in the thinnest era of the corpus. WINDOWS = 10 is what sets the sample size, and it is set against a limit: ten pages is 1.4 MB a day, which is about the most the repository can take.

About

The load-bearing vocabulary of Claude: cluster analysis of GitHub pull requests

Resources

Stars

107 stars

Watchers

2 watching

Forks

Releases

Packages

Contributors

Languages