ARTICLES-All Articles
Longer writeups that tie the experiments together, with more voice than the lab reports.
Testing Needle, a Local Tool-Calling Model A Field Guide to AI Benchmarks GEPA Prompt Optimization on Tiny Local Models CLIR: A Fixed-Vocabulary Intermediate Representation for LLM Code Generation Replicant: Giving AI Agents GPU Access Through Google Colab How Well Do LLMs Follow Length Instructions? Using Claude Code as an AI Research Lab What If Agent Memory Didn't Need a Database? LLMs Can Write Stories but They Can't Remember Them
Testing Needle, a Local Tool-Calling Model
Eleven experiments and roughly 1,800 calls against cactus-needle 2.0.5. The accuracy number turned out to matter much less than which errors it makes, and where.
needletool-callinglocal-modelcactus-needledeterminismevaluationexperiments