Benchmarks
-
On-device LLM leaderboard for boards that fit in 8 GB: tok/s, tok/J, and thermals on a real Jetson. Same prompts and load generator across llama.cpp and Ollama; Pi, phones, and Mac still cooking.0
-
Benchmark evaluating whether language models can generate and edit structured ASCII diagrams. 80 tasks across 4 categories, with public example tasks available on Hugging Face.4
-
Real-phone Android agent benchmark: runs everyday tasks on a live device and scores task success plus phone cost (battery, thermals, $$, steps) per run.0