AI Token Saver aiproxy · by Universally Thinking

Strip bulky context. Spend fewer tokens.

A local proxy between Cursor / Claude Code and the model APIs. It removes duplicate or oversized context before the request goes upstream — then shows you the savings live.

Get in touch View on GitHub

Ops board

Denser KPIs — lifetime tokens and dollars saved, spend velocity, strip rate, noise traffic, and model mix.

aiproxy ops board dark theme with lifetime savings KPIs
Ops board · dark
aiproxy ops board light theme with lifetime and window KPIs
Ops board · light

The dashboard

Full desktop screens — dark and light — for token throughput, disposition, traffic mix, and recent requests.

AI Token Saver dark dashboard showing tokens over time, save rate, and traffic mix
Dashboard · dark
AI Token Saver light dashboard with charts and recent requests
Dashboard · light

What it does

Token Saver runs on your machine. Editor traffic hits the proxy first; bulky or repeated context is stripped, the cleaned request is forwarded, and a dashboard tracks tokens saved, spend velocity, and model mix.

Fewer tokens upstream

Duplicate and oversized payloads are trimmed before they reach the model — often saving more than half of requested tokens.

Live savings dashboard

Watch requested vs saved vs forwarded tokens, save rate, traffic mix, and recent requests in dark or light UI.

Ops board

A denser v2 board for lifetime totals, $/hr velocity, strip rate, noise traffic, and model mix at a glance.

Connect Cursor or Claude Code

Modes match how you already talk to models — MITM for default Cursor, reverse / OpenAI / Anthropic base URLs for BYOK and Claude Code.

Questions or access?

AI Token Saver is a Universally Thinking project. Clone it, run it locally, or reach out for setup help.

Package aiproxy

GitHub benjamin@universallythinking.com