Google introduced Gemini 3.1 Flash-Lite in preview through developer tools and Vertex AI. The announcement emphasizes speed and processing costs, with examples including translation, moderation, interface generation, and other operations repeated at scale.

Context

The relevant cost is a correctly completed request after retries and review. A small model can be evaluated on representative examples, with difficult cases routed to a stronger system or a person. That makes the trade-off measurable: savings should survive the full workflow rather than appear only on the token-price line.

Sources & authors

  1. Gemini 3.1 Flash-Lite: Built for intelligence at scale
    Google DeepMind · March 3, 2026