Personal knowledge base
Save a link, a passage, or a PDF. An agent reads it, writes down what it claims, and folds that into a wiki that already knows what you saved before.
One topic, three sources
Merged into one page. Where they disagree, both stay.
Saved
PaperLLM.int8(): 8-bit matrix multiplication
We find that a small number of feature dimensions — fewer than 0.1% of channels — carry activation magnitudes up to 20× the median, and that quantising these uniformly is what destroys accuracy at scale.
NoteQuantisation, in practice
For anything above 7B, 4-bit weight-only quantisation is effectively free. I have run it across a dozen models and never seen a benchmark move more than a point.
PDFReasoning under 4-bit weights
Aggregate benchmarks conceal the cost. On multi-step arithmetic and chain-of-thought tasks, 4-bit models lose 9–14 points, while single-hop retrieval is unchanged.
Wiki page
Fewer than 0.1% of channels carry outlier activations, up to 20× the median.
“fewer than 0.1% of channels — carry activation magnitudes up to 20× the median”
LLM.int8(): 8-bit matrix multiplication
Quantising those outlier channels uniformly is the dominant source of accuracy loss.
“quantising these uniformly is what destroys accuracy at scale”
LLM.int8(): 8-bit matrix multiplication
4-bit weight-only quantisation is near-lossless above 7B parameters.disputed
“4-bit weight-only quantisation is effectively free”
Quantisation, in practice
4-bit costs 9–14 points on multi-step reasoning, which aggregate benchmarks hide.disputed
“4-bit models lose 9–14 points, while single-hop retrieval is unchanged”
Reasoning under 4-bit weights
The note and the PDF contradict each other. Nothing here decides which is right — both are on the page, marked, with the passage each came from.
And then ask it
Answered only from what this workspace compiled, each sentence carrying the claim it rests on. Conversations are saved, so you can pick one up later.
Why does quantisation hurt accuracy unevenly?
A small number of channels carry outlier activations, and quantising them uniformly is what dominates the errorcitation 1[1] — which is why aggregate benchmarks can look flat while multi-step reasoning dropscitation 2[2].
How it is arranged
Your account
Thesis41 captures · 12 pages
Its own wiki, graph, gaps and conversations.
Work reading18 captures · 6 pages
Its own wiki, graph, gaps and conversations.
Nothing crosses between them. A question asked in one is answered only from what that one has read.
Not a search box
You can ask it things. It answers from what it already compiled, rather than re-reading your library at question time.
Not a folder
Nothing to file, tag or tidy. Pages find their own place and link themselves.
Not a summariser
Every sentence carries the passage it came from. When two sources disagree the page keeps both — deciding for you is how a summary becomes a rumour.
Nothing to configure
Your first page compiles about a minute after your first save.