Benchmarks

  • On-device LLM leaderboard for boards that fit in 8 GB: tok/s, tok/J, and thermals on a real Jetson. Same prompts and load generator across llama.cpp and Ollama; Pi, phones, and Mac still cooking.
    ★ 0
  • Benchmark evaluating whether language models can generate and edit structured ASCII diagrams. 80 tasks across 4 categories, with public example tasks available on Hugging Face.
    ★ 4
  • Real-phone Android agent benchmark: runs everyday tasks on a live device and scores task success plus phone cost (battery, thermals, $$, steps) per run.
    ★ 0