Blog posts

All tags
Jetson
Benchmark
Local LLM
Edge AI
llama.cpp
Ollama
Leaderboard
NVIDIA Jetson
mDNS
AWDL
Networking
Distributed Systems
LLM Inference
1-bit LLM
Cluster Setup
CUDA
Energy Efficiency
GRPO
Reinforcement Learning
LLM Training
Apple Silicon
ML Infrastructure
Raspberry Pi
Python
Edge Compute
Distributed Training
Thunderbolt

No posts match your filters.

2026

1 bit LLMs: The Bonsai Family Story

35 minute read

5 Bonsai-family 1–1.58bit LLMs benchmarked across 4 power modes on Jetson Orin Nano Super 8GB. 25W sweet spot: 47–48% more tok/s than 15W, best output tok/J for all sub-4B models. Read more

◉ — ♡

Clustering 3 Jetson Orin Nano Super

14 minute read

Build a 3-node Jetson Orin Nano Super 8GB cluster with active cooling. Real numbers: ~759 Mbps per link (gigabit), peak 58.3°C across all 3 nodes under full 18-core sustained load, zero throttling at 1728 MHz throughout. Read more

◉ — ♡

smoltorrent: Distributing ML Checkpoints Across a Pi Cluster

24 minute read

A 942 MB checkpoint. Four Raspberry Pis. ~1.5 min gather. No single point of failure. A deep dive into smoltorrent - a distributed checkpoint sharding system built over raw TCP with replication, SHA-256 integrity verification, mDNS discovery, and Prometheus monitoring. Read more

◉ — ♡

Clustering 4 Raspberry Pis 4B

14 minute read

Build a 4-node Raspberry Pi 4B cluster with UCTRONICS enclosure, PoE+ hats, and TP-Link LS110P PoE switch. Real numbers: 94.4 Mbps per link (100 Mbps switch ceiling), 62.3°C under full 16-core load, zero throttling at 1800 MHz throughout. Read more

◉ — ♡

Mac Minis Thunderbolt Cluster Setup Guide

9 minute read

Wire Mac minis into a high-bandwidth local Thunderbolt cluster for distributed training and inference with zero cloud egress cost, low latency, and direct control over cluster networking. Read more

◉ — ♡