Posts by Tags

1-bit LLM

1 bit LLMs: The Bonsai Family Story

35 minute read

5 Bonsai-family 1–1.58bit LLMs benchmarked across 4 power modes on Jetson Orin Nano Super 8GB. 25W sweet spot: 47–48% more tok/s than 15W, best output tok/J for all sub-4B models. Read more

◉ — ♡

AWDL

Apple Silicon

Mac Minis Thunderbolt Cluster Setup Guide

9 minute read

Wire Mac minis into a high-bandwidth local Thunderbolt cluster for distributed training and inference with zero cloud egress cost, low latency, and direct control over cluster networking. Read more

◉ — ♡

Benchmark

1 bit LLMs: The Bonsai Family Story

35 minute read

5 Bonsai-family 1–1.58bit LLMs benchmarked across 4 power modes on Jetson Orin Nano Super 8GB. 25W sweet spot: 47–48% more tok/s than 15W, best output tok/J for all sub-4B models. Read more

◉ — ♡

CUDA

Cluster Setup

Clustering 3 Jetson Orin Nano Super

14 minute read

Build a 3-node Jetson Orin Nano Super 8GB cluster with active cooling. Real numbers: ~759 Mbps per link (gigabit), peak 58.3°C across all 3 nodes under full 18-core sustained load, zero throttling at 1728 MHz throughout. Read more

◉ — ♡

Clustering 4 Raspberry Pis 4B

14 minute read

Build a 4-node Raspberry Pi 4B cluster with UCTRONICS enclosure, PoE+ hats, and TP-Link LS110P PoE switch. Real numbers: 94.4 Mbps per link (100 Mbps switch ceiling), 62.3°C under full 16-core load, zero throttling at 1800 MHz throughout. Read more

◉ — ♡

Mac Minis Thunderbolt Cluster Setup Guide

9 minute read

Wire Mac minis into a high-bandwidth local Thunderbolt cluster for distributed training and inference with zero cloud egress cost, low latency, and direct control over cluster networking. Read more

◉ — ♡

Distributed Systems

smoltorrent: Distributing ML Checkpoints Across a Pi Cluster

24 minute read

A 942 MB checkpoint. Four Raspberry Pis. ~1.5 min gather. No single point of failure. A deep dive into smoltorrent - a distributed checkpoint sharding system built over raw TCP with replication, SHA-256 integrity verification, mDNS discovery, and Prometheus monitoring. Read more

◉ — ♡

Distributed Training

Mac Minis Thunderbolt Cluster Setup Guide

9 minute read

Wire Mac minis into a high-bandwidth local Thunderbolt cluster for distributed training and inference with zero cloud egress cost, low latency, and direct control over cluster networking. Read more

◉ — ♡

Edge AI

Clustering 3 Jetson Orin Nano Super

14 minute read

Build a 3-node Jetson Orin Nano Super 8GB cluster with active cooling. Real numbers: ~759 Mbps per link (gigabit), peak 58.3°C across all 3 nodes under full 18-core sustained load, zero throttling at 1728 MHz throughout. Read more

◉ — ♡

Edge Compute

Clustering 4 Raspberry Pis 4B

14 minute read

Build a 4-node Raspberry Pi 4B cluster with UCTRONICS enclosure, PoE+ hats, and TP-Link LS110P PoE switch. Real numbers: 94.4 Mbps per link (100 Mbps switch ceiling), 62.3°C under full 16-core load, zero throttling at 1800 MHz throughout. Read more

◉ — ♡

Energy Efficiency

GRPO

Jetson

1 bit LLMs: The Bonsai Family Story

35 minute read

5 Bonsai-family 1–1.58bit LLMs benchmarked across 4 power modes on Jetson Orin Nano Super 8GB. 25W sweet spot: 47–48% more tok/s than 15W, best output tok/J for all sub-4B models. Read more

◉ — ♡

Clustering 3 Jetson Orin Nano Super

14 minute read

Build a 3-node Jetson Orin Nano Super 8GB cluster with active cooling. Real numbers: ~759 Mbps per link (gigabit), peak 58.3°C across all 3 nodes under full 18-core sustained load, zero throttling at 1728 MHz throughout. Read more

◉ — ♡

LLM Inference

1 bit LLMs: The Bonsai Family Story

35 minute read

5 Bonsai-family 1–1.58bit LLMs benchmarked across 4 power modes on Jetson Orin Nano Super 8GB. 25W sweet spot: 47–48% more tok/s than 15W, best output tok/J for all sub-4B models. Read more

◉ — ♡

LLM Training

Leaderboard

Local LLM

ML Infrastructure

smoltorrent: Distributing ML Checkpoints Across a Pi Cluster

24 minute read

A 942 MB checkpoint. Four Raspberry Pis. ~1.5 min gather. No single point of failure. A deep dive into smoltorrent - a distributed checkpoint sharding system built over raw TCP with replication, SHA-256 integrity verification, mDNS discovery, and Prometheus monitoring. Read more

◉ — ♡

NVIDIA Jetson

Networking

Clustering 3 Jetson Orin Nano Super

14 minute read

Build a 3-node Jetson Orin Nano Super 8GB cluster with active cooling. Real numbers: ~759 Mbps per link (gigabit), peak 58.3°C across all 3 nodes under full 18-core sustained load, zero throttling at 1728 MHz throughout. Read more

◉ — ♡

Clustering 4 Raspberry Pis 4B

14 minute read

Build a 4-node Raspberry Pi 4B cluster with UCTRONICS enclosure, PoE+ hats, and TP-Link LS110P PoE switch. Real numbers: 94.4 Mbps per link (100 Mbps switch ceiling), 62.3°C under full 16-core load, zero throttling at 1800 MHz throughout. Read more

◉ — ♡

Ollama

Python

smoltorrent: Distributing ML Checkpoints Across a Pi Cluster

24 minute read

A 942 MB checkpoint. Four Raspberry Pis. ~1.5 min gather. No single point of failure. A deep dive into smoltorrent - a distributed checkpoint sharding system built over raw TCP with replication, SHA-256 integrity verification, mDNS discovery, and Prometheus monitoring. Read more

◉ — ♡

Raspberry Pi

smoltorrent: Distributing ML Checkpoints Across a Pi Cluster

24 minute read

A 942 MB checkpoint. Four Raspberry Pis. ~1.5 min gather. No single point of failure. A deep dive into smoltorrent - a distributed checkpoint sharding system built over raw TCP with replication, SHA-256 integrity verification, mDNS discovery, and Prometheus monitoring. Read more

◉ — ♡

Clustering 4 Raspberry Pis 4B

14 minute read

Build a 4-node Raspberry Pi 4B cluster with UCTRONICS enclosure, PoE+ hats, and TP-Link LS110P PoE switch. Real numbers: 94.4 Mbps per link (100 Mbps switch ceiling), 62.3°C under full 16-core load, zero throttling at 1800 MHz throughout. Read more

◉ — ♡

Reinforcement Learning

Thunderbolt

Mac Minis Thunderbolt Cluster Setup Guide

9 minute read

Wire Mac minis into a high-bandwidth local Thunderbolt cluster for distributed training and inference with zero cloud egress cost, low latency, and direct control over cluster networking. Read more

◉ — ♡

llama.cpp

mDNS