Gemini 2.5 Flash Lite
Google logo

Gemini 2.5 Flash Lite

gemini-2.5-flash-litellms.txt
Google
Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.

Pricing

PricingCache ReadInput AudioInput Audio Cached Web SearchCache Storage
$0.1$0.4
$0.01/M tokens$0.3/M tokens$0.03/M tokens$0.035/request$1/h/M tokens

Input Modalities

  • Text
  • Vision
  • Audio
  • Video
  • PDF

Output Modalities

  • Text

Context length

  • 1.05M tokens

Max output

  • 65.5K tokens

Capabilities

  • Thinking
  • Streaming
  • Tool calling
  • Web search
  • URL context
  • Code interpreter
  • Computer use
  • File search
  • Memory tool
  • Structured outputs
  • Citations
  • Prompt caching
  • Background mode
  • Server-side sessions

Providers

VertexAI gemini-2.5-flash-lite
Pricing$0.1$0.4
Cache Read$0.01/M tokens
Input Audio$0.3/M tokens
Input Audio Cached $0.03/M tokens
Web Search$0.035/request
Cache Storage$1/h/M tokens
Context1M
Max output65K
Latency2.1S
Throughput50.4TPS
Uptime
99.93% uptime 2 days ago
99.47% uptime yesterday
100.00% uptime today
Google AI Studio gemini-2.5-flash-lite
Pricing$0.1$0.4
Cache Read$0.01/M tokens
Input Audio$0.3/M tokens
Input Audio Cached $0.03/M tokens
Web Search$0.035/request
Cache Storage$1/h/M tokens
Context1M
Max output65K
Latency0.9S
Throughput26.6TPS
Uptime
96.50% uptime 2 days ago
95.59% uptime yesterday
95.20% uptime today

Performance for gemini-2.5-flash-lite

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="gemini-2.5-flash-lite",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is Gemini 2.5 Flash Lite?

Gemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.