This product announcement is available in English.
Trio-Spark v1.1: See the scene. Choose the next move.
A hand signal changes. A person enters the frame. A situation crosses from one allowed action to another. The next move is often visible before it can be written down.
Trio-Spark v1.1 can see that moment and choose what comes next. Bring text, one image, or a short sampled window to one model, one API key, and one wallet. Define the moves your application allows. Spark returns one choice and a probability for every option, without writing an answer first.
Bring the world at hand
Text, images, and sampled camera windows now meet in the same decision product. An existing loop can add what is visible while keeping the same bounded choice contract: the situation and allowed moves go in; one move and the full distribution come back.
In the Playground, start with an uploaded image, a local video, or your camera. The gesture demo makes the loop tangible: show the scene, define the gestures that matter, and watch Spark choose between them.
Choices stay in your application
Spark chooses only from the two to eight actions you supply. Your application can include options such as wait, ask for review, or a small set of domain actions, then decide how to use the returned probabilities. The result is ready for code to inspect rather than prose for code to interpret.
Try a visual decision
The new gesture demo shows the loop end to end. Choose the visual input mode, allow camera access or upload media, and define the gesture choices. The result includes the selected choice, the full distribution, and usage.
Open the gesture demo or read the API docs for image and sampled-video request examples.
Follow a scene in motion
A short sampled window can show how a scene changes. Here, Spark follows movement through Shibuya Crossing and chooses which road users are visibly moving through the center.
Measured against Clef-Flash
Spatial judgment is a useful test of a situated model: not just what is visible, but how things relate. On our 95-question GQA spatial subset, Spark answered 81 correctly versus Clef-Flash’s 73 — an 8.4 percentage-point lead. Spark also scored higher on the counting, scene-semantics, and attribute subsets below.
| Task | Questions | Spark | Clef-Flash |
|---|---|---|---|
| Static yes/no | 154 | 93.5% (144/154) | 94.2% (145/154) |
| Color | 38 | 100.0% (38/38) | 100.0% (38/38) |
| Counting | 48 | 91.7% (44/48) | 89.6% (43/48) |
| Scene semantics | 80 | 86.3% (69/80) | 82.5% (66/80) |
| Static set · total | 320 | 92.2% (295/320) | 91.3% (292/320) |
| GQA object existence | 71 | 77.5% (55/71) | 80.3% (57/71) |
| GQA attributes | 228 | 77.2% (176/228) | 75.4% (172/228) |
| GQA spatial relations | 95 | 85.3% (81/95) | 76.8% (73/95) |
| GQA global (small diagnostic) | 6 | 83.3% (5/6) | 83.3% (5/6) |
| GQA subset · total | 400 | 79.3% (317/400) | 76.8% (307/400) |
| Fresh video questions | 120 | 54.2% (65/120) | 56.7% (68/120) |
These are results from a frozen v1.1 development snapshot on October 3, before subsequent serving updates; they are not a new evaluation of the current production revision. Static and GQA results use opened subsets, with matching questions and images, rather than full benchmark test sets. Existing Spark runs were paired with later Clef-Flash calls. One Spark timeout and one Clef authorization failure remain in the denominators. Total rows summarize the preceding subgroups.
The comparison has tradeoffs: Clef-Flash scored higher on object existence and the fresh video set, and made fewer high-probability errors across the static and GQA sets (maximum option probability ≥ 0.8; these probabilities are not calibrated). The video difference was inconclusive at this sample size. These results support specific strengths, not a claim of overall superiority.
A short path from pixels to a decision
On a T4, our eight-clip gesture diagnostic measured a median 669 ms of origin processing for a single image and 977 ms for four sampled frames. Measured client time, including the two HTTPS calls in that test, was 1.26 s and 1.63 s respectively. A still gesture can use the latest frame; motion can use a short sequence.
Small development diagnostic, not a latency guarantee or gesture-accuracy benchmark. Camera capture, browser rendering, network conditions, and queueing affect the live experience. Clef used a different hosted path, so we do not infer a hardware-normalized speed comparison. Download the counts and measurement notes.
Simple usage-based pricing
Trio-Spark v1.1 costs $0.042 per million billed input tokens for text and visual requests. The Playground includes 100 free decisions, with no card required. API responses report input-token usage so teams can measure each workflow directly.
Developer details
Use trio-spark-v1.1 for text, one JPEG or PNG image, or two to four timestamped frames from a window of up to 30 seconds. The API docs cover the request schema, media bounds, context limits, errors, and idempotent retries. Our published text benchmarks and recorded game demos remain clearly labeled v1.0 results.
Build the next move
Start with a narrow decision your application can verify, define the allowed actions, and bring the state as text, one image, or a short sampled window.
Try Trio-Spark v1.1