Evidence at a glance
The mechanism in one line
Compress the visual or contextual input before the main reasoning path.
Route or verify the expensive step instead of repeating the full path.
Translate the mechanism into a bounded deployment or evaluation check.
Keep the Agent, Make the Model Replaceable
Together AI released Together Link on October 5, 2026. The beta CLI is described by the company as free and MIT-licensed, and is aimed at developers already using Claude Code, Codex, OpenCode, Pi, or desktop apps. It does not provide a new coding agent; it connects those tools to models hosted by Together AI.
The design addresses a practical tension: teams may prefer an agent's workflow without wanting every task to go to the same expensive model. Together Link's proposition is to keep the existing tool harness while separating model choice from agent preference. The change is in the request path and model bill, not in the basic way developers work with their agents.
A Lightweight Setup, but Inference Stays in the Cloud
The installation entry point is short: the documentation provides a single curl command, but users still need a Together API key and the target tool already installed. Developers can launch sessions with commands such as `togetherlink claude` or `togetherlink codex`. The supported entry points include Claude Code, Claude Desktop, Codex, ChatGPT Desktop, OpenCode, and Pi, though their configuration methods are not identical.
The key architectural choice is that requests pass through Together's cloud gateway; the official documentation says no local proxy or daemon runs. Terminal tools use temporary configuration that is removed when a session ends, while Claude Desktop and ChatGPT Desktop use separate profiles that can be switched back. Here, “open models” means models hosted by Together and called through its API, not models running on a developer's machine. Setup gets shorter, but inference still depends on the provider, connectivity, and a Together account.
Auto Routing Drives the Savings Case—and Raises Questions
Together Link defaults to a virtual Auto model, while also allowing users to pin a model. The launch announcement describes routing based on the first task in a session: quick fixes go to faster, lower-cost options, while harder problems get stronger capabilities. It says each session is routed only once, preserving prompt caching. With an Anthropic API key, Claude Code and Claude Desktop can also route between Opus 5.5 and GLM 5.3; Opus usage is billed to the user's Anthropic account. Other supported tools use Together models.
However, public materials disagree about routing frequency. The launch announcement says the choice is made once per session, while the current documentation says the cloud gateway classifies and routes each request. Those approaches can affect caching, cost estimates, and whether the same task behaves consistently, so the difference is more than wording. The product is still in beta; teams should verify how the current version behaves before relying on Auto rather than treating either description as a stable contract.
Prices Show the Cost Structure, Not the Savings
The documentation lists prices per million tokens: GLM 5.3 costs $1.40 for input and $4.40 for output, while DeepSeek V4.1 Flash costs $0.30 and $1.20. Kimi K3 is listed in the documentation at $2.70 for input and $13.50 for output, but the product page lists $3 and $15. The public materials do not explain whether the difference reflects a price change, a page update, or something else; actual costs should be checked against runtime pricing and bills.
Together claims savings of more than 50%, and its product page gives a range of 50% to 80%. The available materials do not provide independently verifiable task samples, quality comparisons, or a calculation method. A lower token price does not guarantee a lower total cost: comparisons also depend on input and output tokens, retries, which model the router selects, and whether Claude usage adds charges to an Anthropic account. Session receipts and `togetherlink usage --last 7d` provide ways to inspect spending, but they do not replace comparisons against task quality and a cost baseline.
Treat It as a Testable Integration Layer, Not a Savings Guarantee
For technical leaders, the main question is not whether Together Link can be installed with one command, but whether it makes model substitution a controllable engineering variable. Start with a representative set of tasks in the existing agent workflow, recording the model used, token consumption, bill, and completion outcome, then compare those results with a fixed-model baseline. If using Auto, check the selected routes and any Anthropic-billed usage separately so spending across accounts is not folded into a misleading savings figure.
Convenience has limits: this is not a local inference solution, support is limited to macOS and Linux, and the product is in beta, with commands, routing, and model availability subject to change. For production workflows, the right response is neither to assume the advertised savings nor to reject the tool outright because its routing descriptions differ. Instead, make price discrepancies, routing behavior, and fallback procedures part of a small pilot's acceptance criteria. The connector earns a place in a team's toolchain only when the team can reproduce both the bill and the quality results on its own tasks.