Trio-Spark v1.0 · Now live

Introducing Trio-Spark v1.0: A one-pass decision API for GUI agents

Product update: Trio-Spark v1.1 now brings one image or a short sampled visual window to the same bounded-choice API. Read the v1.1 visual launch update. The benchmarks and recorded runs below remain v1.0 results.

Most agent loops do not need another paragraph. They need a next move.

Open the export menu. Escalate the refund. Pause the machine. Place the falling block. These are small decisions, but they are where an agent meets the world—and where latency, cost, and vague output compound fast.

Trio-Spark v1.0 is built for that moment. Send the situation and two to eight actions your software is allowed to take. Spark returns one choice, a probability for every option, and usage your code can inspect. No essay. No action invented outside your list. No text to parse before the loop can continue.

A decision model for the world at hand

Spark is the first public model in our Situated World Models program. The idea is simple: useful intelligence should be shaped around the environment it serves. A browser, a game, a factory floor, and a restaurant do not share the same state or the same possible moves. The application defines both; Spark supplies fast judgment between them.

That makes Spark a natural decision layer for GUI agents and interactive systems. Your agent observes the current state, enumerates valid actions, calls Spark, and executes or reviews the result. Then it observes again. The model stays inside the real loop instead of sitting beside it as a chatbot.

Built to decide, not to write

Generative models spend time producing a string and leave your application to interpret it. Spark scores the actions together in a single model pass and generates zero output tokens. Its API response is already the thing your program needs: the selected action ID, the full probability distribution, and the model version that made the call.

Because the action space comes from your code, you keep control of what can happen next. Add a wait action. Remove a destructive action before calling. Route low-confidence cases to a person. Replay the same state against a new model version. Spark turns model judgment into a component you can measure.

Fast enough for a live agent loop

We measured a controlled 64-step agent loop on T4-class GPUs. After the first decision established the session, subsequent calls completed in 83 ms median, compared with a 690 ms full-context production baseline—a median 8.3× speedup. All 64 selected actions matched the baseline.

Continuous-session latency on a controlled T4 loop
MeasurementResult
Steady-state median83 ms
Median speedup8.3×
Decision agreement64 / 64

This gain applies to continuous sessions with substantial repeated context. A first call or an unrelated one-off request still pays the full prompt cost. Network distance and input length also affect end-to-end latency.

The first release, measured

We ran Trio-Spark v1.0 through the production T4 API across full public task sets. It completed every example in each result below.

Trio-Spark v1.0 benchmark figure showing 92.89 percent on SST-2, 88.46 percent on AG News, and 80.58 percent on TweetEval Offensive, all with complete valid-output coverage
Trio-Spark v1.0 release snapshot · frozen evaluations through the production T4 API · September 30, 2026
Complete public-task evaluations
EvaluationProtocolValid outputsAccuracy
SST-2Binary sentiment872 / 87292.89%
AG NewsFour-way topic7,600 / 7,60088.46%
TweetEval offensiveBinary classification860 / 86080.58%

Every response includes probabilities and a decision status. Applications can act on clear calls and send ambiguous ones down a review path, without changing the model interface.

Two games, one decision loop

Falling Blocks: each move creates the next state

Each piece starts with a board state and a short list of legal placements. In this 30-second production replay, Trio-Spark v1.0 makes 16 decisions and clears seven lines. Every selected placement changes the board sent in the next request.

Production replay · 30 seconds · 16 decisions · 7 lines cleared

2048: the same loop, with different choices

Here the state is a 4×4 grid and the allowed actions are legal directions. This 39-second production replay follows 26 consecutive decisions as the game applies Spark's choice and returns the new board. The same request, choice, and next-state loop now runs against a different world.

Production replay · 39 seconds · 26 decisions

Games make the feedback easy to see. In a browser or operations workflow, the state and action names change; the loop stays the same: observe, enumerate allowed actions, decide, act, and observe again.

Cheap enough to stay in the loop

Trio-Spark v1.0 costs $0.042 per million billed input tokens, with free output and no subscription. New accounts receive 100 free trial attempts in the Playground. After that, the minimum top-up is $5 and usage remains pay as you go.

The service is live today. Start with the game loop, replace the example with one of your own decisions, then create an API key when you are ready to connect your application.

Results are for Trio-Spark v1.0 on the stated evaluations and hardware. The latency result is a controlled 64-step continuous-session loop after the first call, rather than an SLA for unrelated requests. Test decisions and confidence thresholds on your own workload before automating consequential actions.