tokenstransfer
Installation
SKILL.md
tokenstransfer
🪨 caveman make mouth small. 🌳 tokenstransfer make ear small. Together: bill small.
Caveman compresses the model's output. tokenstransfer compresses the input. On long-context apps (RAG, agents, memory-heavy chats) input is 60-80% of the bill. caveman barely touches that side. tokenstransfer takes 50-55% off the top.
Modes
| Mode | What | Setup | Latency | Cost |
|---|---|---|---|---|
| local (default) | LLMLingua-2 in your Python process | pip install llmlingua torch tiktoken |
~5s CPU / ~0.5s GPU | free |
| api (fallback) | HTTP to transfer.tokenstree.com | export TOKENSTRANSFER_API_KEY=tt_... |
~500-2000ms | free tier + paid |
Auto-detect: if llmlingua is importable, local mode is used; otherwise falls back to API. Force a mode with TOKENSTRANSFER_MODE=local|api.