tokenstransfer

Installation
SKILL.md

tokenstransfer

🪨 caveman make mouth small. 🌳 tokenstransfer make ear small. Together: bill small.

Caveman compresses the model's output. tokenstransfer compresses the input. On long-context apps (RAG, agents, memory-heavy chats) input is 60-80% of the bill. caveman barely touches that side. tokenstransfer takes 50-55% off the top.

Modes

Mode What Setup Latency Cost
local (default) LLMLingua-2 in your Python process pip install llmlingua torch tiktoken ~5s CPU / ~0.5s GPU free
api (fallback) HTTP to transfer.tokenstree.com export TOKENSTRANSFER_API_KEY=tt_... ~500-2000ms free tier + paid

Auto-detect: if llmlingua is importable, local mode is used; otherwise falls back to API. Force a mode with TOKENSTRANSFER_MODE=local|api.

Trigger

Installs
1
First Seen
Jun 1, 2026
tokenstransfer — vfalbor/caveman_tokenstransfer.2.0