GoKiteAI announced MarathonBuild, adaptive inference infrastructure for long-running agents that lets each request use a completion window from immediate to anytime, with the source stating costs can be up to about 65% lower when users can wait.
The update adds a compute layer that can make agent inference more programmable and potentially cheaper. By letting jobs wait when they do not need an instant response, the scheduler can use available capacity more efficiently while keeping the same models and API. That can broaden Kite's utility for agent workflows that run across many steps, retries, or extended tasks.
The practical relevance still depends on whether developers use MarathonBuild in meaningful volume and whether that activity translates into direct demand for KITE. The launch confirms availability of the infrastructure, but it does not by itself show adoption or a demonstrated token link.
We are introducing , incubated by : adaptive inference infrastructure for long-running agents.
Intelligence keeps getting cheaper, but the monthly bill does not follow. Agents live in that gap: they run long jobs, retry, and plan across many steps, and…