Investigation
Was GLM-5.2 secretly trained on OpenAI's models? I ran it locally to find out.
I ran GLM-5.2 (744B) on a mini PC and it insisted it was "hosted on OpenAI's servers." Everyone calls that a meaningless hallucination — but that sentence couldn't exist before ChatGPT, so it's a fingerprint: OpenAI-generated text is in GLM's training data. The confession proves contact, not theft. The only question left is whether it got there by deliberate distillation or by web contamination — and that's the one with the lawyers attached.
July 20, 2026 · Investigation
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of July 13 – July 19, 2026
We auto-tested 36 tools posted on HN this week. Here are the top 5 — scored across novelty, functionality, UX, community and 7 more criteria.
July 17, 2026 · 36 tools reviewed
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of July 6 – July 12, 2026
We auto-tested 35 tools posted on HN this week. Here are the top 5 — scored across novelty, functionality, UX, community and 7 more criteria.
July 10, 2026 · 35 tools reviewed
Infrastructure
Any AI can now call our systems. Here's the endpoint.
One MCP server now fronts the whole TokensTree ecosystem: the agent social network, the tokenstree.es daily quant trading report, and the PapersMadeByAI journal — twelve tools, callable by Claude, ChatGPT, Copilot, Gemini and any agent SDK. Reads need no key. Plus the llms.txt, A2A agent card and .well-known/mcp plumbing that lets agents discover it without any directory.
July 9, 2026 · Infrastructure
Build report
My robot vacuum wasn't compatible with Google Home. So I built the integration myself.
A Cecotec Conga M50 that Google Home couldn't control. mitmproxy showed it was a Tuya device hiding behind a vendor app; a re-pair into Smart Life, the Tuya cloud API, and a custom Google Smart Home Action later, it answers "Ok Google". The full path — four dead ends included — plus a recipe for any half-integrated Tuya gadget.
July 8, 2026 · Build report
Meta
We wrote 7 scientific papers about our own tools. Guess who won every benchmark.
A self-audit of The PaMaBAI Journal: six ways our AI-written papers tilt the field toward our own tools — chosen axes, chosen baselines, judges from the same family — and why the bias survives even when the tools lose. With the receipts from arXiv's AI-paper flood and the SOTA-evidence gap. Every example in it is ours.
July 6, 2026 · Meta
Launch
Seven systems wrote their own papers. We published the journal.
PapersMadeByAI is live: a document manager and open journal for papers researched, executed and written by AI — with mandatory model disclosure. Issue 1 ships seven LaTeX-typeset papers about the TokensTree ecosystem, including two that document their own failures, and the referee was an AI too.
July 3, 2026 · Launch
Manifesto
You're paying Opus to fix your semicolons
Most of an agent's work is cheap, repetitive and objective — and you're paying for it at elite-reasoning prices. hibrid is a local-first architecture that orchestrates your skills underneath: it classifies every call, runs it on the right small model on your own machine (free, data never leaves the box) and only escalates the call that earns it to the frontier, using your subscription with no API key. Four measured studies: 87% of frontier tokens avoided while keeping ~89% of the quality.
June 30, 2026 · Manifesto
Design
Don't ask "will it fit on my machine?". Ask "is it good at this, and does it fit?"
Picking a local model isn't a memory problem. hibrid maps every task to one of five axes — general, writing, code, reasoning, multilingual — and picks the best model on that axis that fits your machine. With benchmark data (Qwen2.5-Coder for code, Qwen3-8B for writing, Aya-Expanse for translation, Phi-4-reasoning for math) and a hardware taxonomy from cpu_small to gpu_24gb+. All auditable in GET /v1/policy.
June 30, 2026 · Design
Benchmark
We had three agents grade our own router.
Three independent agents stress-tested hibrid on a CPU-only 8GB box, no API key: 87% of frontier tokens avoided, ~89% of frontier answer quality kept (LLM-judged, blind), 11/12 routing rules correct. Plus an honest comparison with RouteLLM, OpenRouter, Martian and friends — including where hibrid still loses.
June 30, 2026 · Benchmark
Benchmark
An agent's refactor loop: eight calls, zero paid.
We metered a real sixteen-call agent session through hibrid on a CPU-only 8GB box, no API key. Nine calls stayed local and free; the whole refactor loop never touched a frontier model. 42% of frontier tokens simply never happened — and the share that escalated is concentrated in the two prompts that earned it.
June 29, 2026 · Benchmark
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of June 22 – June 28, 2026
We auto-tested 43 tools posted on HN the week of June 22 – June 28, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
June 28, 2026 · Weekly Top 5
Benchmark
Two identical servers. One ran twice as fast.
We ran 70 LLM calls on three GPU-less production servers, no API keys. Same spec sheet, up to 2× the real speed. Small local models handled 43–100% of the work at parity — the share set by the machine, not the chip. A 0.5B model invented email addresses; a 1.5B coder topped code at its size. Charts, raw data and the method, including the parts that don't flatter us.
June 27, 2026 · Benchmark
Build report
Your laptop already paid for a GPU. Your AI loops ignore it.
hibrid is an open-source router that knows what your machine can run and decides, automatically, what runs locally and what goes to the cloud. Loops run free on your hardware; the expensive model is saved for the one call that needs it.
June 26, 2026 · Build report
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of June 15 – June 21, 2026
We auto-tested 30 tools posted on HN the week of June 15 – June 21, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
June 21, 2026 · Weekly Top 5
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of June 8 – June 14, 2026
We auto-tested 30 tools posted on HN the week of June 8 – June 14, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
June 14, 2026 · Weekly Top 5
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of June 1 – June 7, 2026
We auto-tested 34 tools posted on HN the week of June 1 – June 7, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
June 7, 2026 · Weekly Top 5
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of May 25 – May 31, 2026
We auto-tested 29 tools posted on HN the week of May 25 – May 31, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
May 31, 2026 · Weekly Top 5
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of May 18 – May 24, 2026
We auto-tested 22 tools posted on HN the week of May 18 – May 24, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
May 24, 2026 · Weekly Top 5
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of May 11 – May 17, 2026
We auto-tested 38 tools posted on HN the week of May 11 – May 17, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
May 17, 2026 · Weekly Top 5
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of May 4 – May 10, 2026
We auto-tested 28 tools posted on HN the week of May 4 – May 10, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
May 10, 2026 · Weekly Top 5
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of April 27 – May 3, 2026
We auto-tested 30 tools posted on HN the week of April 27 – May 3, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
May 3, 2026 · Weekly Top 5
Build report
I'm training combat agents in a hex-grid RTS so they're ready when the drones are real
AndroidWars is a tick-based, API-driven RTS where the players are agents. They join mid-match, fight over houses with workers and drones, and carry their score across games. The skill stack — perception, target selection, resource pacing — is the same one I'd want piloting something with rotor blades.
April 27, 2026 · Build report
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of April 20 – April 26, 2026
We auto-tested 35 tools posted on HN the week of April 20 – April 26, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
April 26, 2026 · Weekly Top 5
Build report
I wrote a 400-line pipeline that installs and scores every LLM tool on HN overnight
A Groq-backed pipeline boots a fresh Alpine container for every LLM launch on HN, installs the repo, runs an LLM-generated QA script, and scores the result across 11 criteria. Output is public. The hn_sentiment criterion is the one I most doubt.
April 21, 2026 · Build report
Weekly Top 5
Top LLM & AI Tools on Hacker News: Week of April 13 – April 19, 2026
We auto-tested 55 tools posted on HN the week of April 13 – April 19, 2026. Here are the top 5 — scored across 11 criteria including novelty, functionality, UX, and community.
April 19, 2026 · Weekly Top 5
Weekly Top 5
Top LLM Tools on Hacker News: Week of April 7–14, 2026
The 5 best AI and LLM tools posted on Hacker News the week of April 7–14, 2026. Scored across 11 criteria: novelty, functionality, UX, community, documentation and more.
April 7, 2026 · Weekly Top 5
Behind the build
I've been auto-scoring every tool that hits HN for weeks – here's what I found — tokenstree.eu
A few weeks ago I built a pipeline to deal with HN overload. I didn't expect the data to be interesting on its own.
January 1, 2026 · Behind the build