gemini-2.5-flash-lite-nothink
GoogleGemini 2.5 Flash-Lite is a balanced model from Google, optimized for applications that require low-latency performance. It retains the practical capabilities of the Gemini 2.5 family, including configurable reasoning based on budget, integration with tools such as grounding via Google Search and code execution, multimodal input support, and an ultra-long context window of up to 1 million tokens, delivering a strong balance between efficiency, functionality, and cost.
Pricing
Input Modalities
- Text
- Vision
- Audio
- Video
Output Modalities
- Text
Context length
- 1.05M tokens
Max output
- 65.5K tokens
Capabilities
- Thinking
- Streaming
- Tool calling
- Web search
- URL context
- Code interpreter
- Computer use
- File search
- Memory tool
- Structured outputs
- Citations
- Prompt caching
- Background mode
- Server-side sessions
Try this model
Python