DeepSeek V4.1 Flash is attracting attention after third-party hosting platforms listed its input token prices at nearly 1,500 times below the company’s usual off-peak rate.
On OpenRouter, the Relace service lists input processing at $0.0001 per million tokens, compared with DeepSeek’s official off-peak rate of $0.15 per million tokens. Open Inference offers a similar rate of approximately $0.00011 per million input tokens.
However, the significant price difference applies specifically to input processing rather than every aspect of using the model. Output costs remain considerably higher, meaning developers may not achieve the same savings when generating lengthy responses.
Third-Party Platforms Offer Lower Input Rates
Relace lists output processing at $0.60 per million tokens, matching DeepSeek’s official off-peak output price. Open Inference lists output at approximately $0.36 per million tokens.
These prices could benefit developers running applications that process large amounts of text, particularly coding assistants and automated agents that repeatedly analyse extensive instructions or documents.
Nevertheless, third-party providers can change their prices, and the actual cost depends on factors such as caching, routing and the volume of generated output.
Developers must also consider performance and reliability before choosing a provider based solely on its advertised rates.
DeepSeek V4.1 Flash Targets Advanced AI Workloads
DeepSeek introduced V4.1 Flash in September 2026, offering native image understanding, a context window of up to one million tokens and capabilities aimed at coding and agent-based tasks.
The model uses a mixture-of-experts architecture with a 552-billion-parameter backbone. It activates approximately eight billion parameters during input processing and 16 billion during output generation, helping improve computational efficiency.
Its architecture also incorporates techniques designed to reduce memory requirements when handling lengthy conversations and repeated prompts.
The latest third-party listings highlight how hosting competition can create substantial differences in AI service costs. However, developers should compare the complete pricing structure, response quality, speed and uptime before deciding whether a cheaper endpoint suits their needs.
For the latest updates, visit and follow The Truth International website (www.thetruthinternational.com) and subscribe to the YouTube Channel.
