EveryGPU
Distributed LLM inference engine
An experiment in joining spare GPUs into an efficient inference pipeline, with instrumentation that shows where every request waits.
- Status
- In progress
- Period
- Sep 2026 - Present
- Built with
- Python
- CUDA
- TCP
- OpenTelemetry
What it does
- Built a distributed inference prototype that passes activation vectors over TCP from one GPU, through a laptop-hosted server, to the next GPU shard.
- Built an evaluation platform that breaks every request into end-to-end latency, per-shard prefill and decode time, vector transport overhead, and server-queue wait time.
- Now optimizing a small model toward 15-20 tokens per second before testing a larger model across more distributed GPU networks.
Why I built it
EveryGPU began with a simple question: if a laptop, Colab, and Kaggle can each offer usable GPU capacity, why can’t they work together to run a model that none could serve alone? I started by exploring how scattered, otherwise-idle GPUs could act as one inference system.
The hard part
The first prototype made clear that adding GPUs is not enough. Each request must be split, scheduled, and carried between machines without letting queueing or network transfer erase the gains from extra compute. Right now, activation vectors travel over TCP through a laptop-hosted server between GPU shards.
What I learned
Distributed inference needs measurement before optimization. The evaluation platform separates queue wait, per-shard prefill and decode, and transport overhead, so I can see where a request actually spends time.
Where it goes next
The immediate target is an efficient small-model system at roughly 15-20 tokens per second. From there, I want to test direct peer-to-peer transport, potentially with QUIC, and scale to larger models and more remote GPUs. The longer-term work is routing requests well across n GPUs and m shards, sustaining useful concurrent throughput, recovering from unreliable nodes, and building a desktop app that gives a participant’s GPU back when they need it.
How it works
Prompts
requests enter the system
Router
splits and schedules work; a laptop server today
activation vectors over TCPGPU shards
spare laptop, Colab, and Kaggle GPUs, each running one model stage
Tokens
generated output
An evaluator records server-queue wait, per-shard prefill and decode, transport overhead, and end-to-end latency. Next: direct peer-to-peer transport, n GPUs by m shards, and recovery from unreliable nodes.