Skip to content

Inference · Python

ProppaTP

CPU-stage prototype

ProppaTP explores prompt preprocessing and tensor-parallel LLM decoding on consumer GPUs. The published project is a CPU-stage prototype; GPU execution, nvFP4 loading, and the prefill/decode split are development targets, not demonstrated throughput results.

github.com/AStarStarship/… →

Specification

LanguagePython (PyTorch)
Target rig1x RTX 5070 Ti (prefill) + 2x RTX 5060 Ti (TP=2 decode)
Target interconnectPCIe 5.0 x16 per card (no NVLink)
Quantization targetnvFP4; model loading is pending

Design & implementation

Planned prefill / decode split

The target architecture assigns prompt encoding to one GPU and tensor-parallel decode to a pair of GPUs. The public checklist still marks this GPU split as unfinished.

Row-parallel over PCIe

The design uses row-parallel linear operations to limit communication over PCIe without NVLink. TP=2 shard execution and GPU all-reduce still need hardware validation.

Cache development

Paged KV caching is a development target. GPU cache behavior and memory use need measurement before the project can claim a throughput improvement.

Target model allocation

The README outlines separate auxiliary and decode models. Their nvFP4 loading, GPU placement, and memory budgets are targets to validate, not a working multi-model deployment demonstrated here.

Related