Publicado el 04 sept 2026 · Confirmamos el 04 sept 2026 que sigue activo
AU$ 15 – AU$ 25 por proyecto
I’m burning through roughly 850 000 000 tokens every day while experimenting with Vibe coding projects that rely on Claude Code + Skills, and the bill is piling up fast. Most of that waste appears when the code calls Claude and, to a lesser extent, Ollama; Qwen 3.8 is in the mix too, but its footprint is smaller. Locally the workflow runs, but it still eats tokens, and on the H200 server I spun up on Vast the same stack becomes painfully slow as well as expensive. Here’s what I need: • A careful audit of the current prompt chains and function calls to pinpoint exactly where Claude Code and Ollama are over-allocating tokens. • Concrete code-level optimisation techniques—prompt refactoring, context window trimming, token-length guards, caching strategies or any other proven methods—that cut the total daily tokens consumed. • Configuration tweaks for both my local machine and the remote H200 server (CUDA, model quantisation, batching, concurrency limits, etc.) so performance improves instead of degrading. • Suggestions for alternative tools, models, or routing logic if replacing parts of the stack would save more tokens than patching them. Deliverables I’d like to see: 1. A brief report that highlights every hotspot you find, shows before-/after-token counts and explains what changed. 2. Updated or annotated source files and configuration snippets ready for me to drop into the current repo. 3. A concise setup guide that walks me through replicating the optimised environment locally and on the Vast server. If your fixes slash token use significantly without slowing things down, that will be my success metric. I’m ready to provide access to logs, sample prompts and any other details you need to get started.
Crea una cuenta gratis para ver el empleo completo y postularte.