New AI Model Balances Power and Cost for Developers

TL;DR: Thinking Machines' new Inkling Small AI model is now on Vercel's AI Gateway. It offers performance similar to larger models at a fraction of the size and cost, making advanced, multimodal AI more accessible for developers.
Key facts
- Category
- AI
- Impact
- High
- Published
- Source
- Vercel Blog
Full summary
Vercel's AI Gateway now offers Inkling Small, a new model delivering high performance at a quarter of the size and compute cost.
Vercel announced that the Inkling Small AI model from Thinking Machines is now available on its AI Gateway platform. According to the Vercel Blog, this new model is designed for efficiency, delivering performance comparable to its larger predecessor but at approximately a quarter of the size. This reduction in scale means it requires significantly less computational power for each task. Inkling Small is a versatile, generalist model capable of natively processing and reasoning over various data types, including images and audio. Its capabilities extend to complex reasoning, tool use, and agentic coding, where an AI can autonomously write or modify code to solve problems, making it a powerful new option for developers building on the Vercel ecosystem.
The core innovation of Inkling Small lies in its balance of performance and resource consumption. By shrinking the model's size without a proportional drop in capability, Thinking Machines addresses a major challenge in the AI industry: the high cost of running large, powerful models. This efficiency is achieved through advanced model architecture and training techniques. A key feature highlighted is its “controllable thinking effort,” which allows developers to adjust the trade-off between response quality and computational cost. In practice, this could mean developers can set the model to a faster, less resource-intensive mode for simple queries, while dialing up the “effort” for more complex tasks that require deeper analysis, providing granular control over application performance and budget.
This development is significant for a wide range of technology professionals. For developers, it provides a cost-effective yet powerful tool for integrating sophisticated, multimodal AI features into their applications directly through Vercel's familiar infrastructure. For CTOs and engineering leaders, Inkling Small offers a path to leverage advanced AI without incurring the steep operational costs and infrastructure overhead typically associated with larger models. Furthermore, the selection reason noted the model’s Zero Data Retention policy. This is a critical feature for security and compliance teams, as it ensures that sensitive data processed by the model is not stored, mitigating privacy risks and helping organizations meet regulatory requirements like GDPR.
The introduction of Inkling Small on a major developer platform like Vercel AI Gateway reflects a broader industry shift away from a singular focus on building ever-larger AI models. The market is now seeing a growing demand for smaller, more efficient, and specialized models that provide a better return on investment for specific business use cases. This trend democratizes access to advanced AI, allowing smaller companies and independent developers to build products that were previously only feasible for large, well-funded organizations. The practical takeaway for businesses is to evaluate these new, efficient models for their AI workloads, as they can unlock new product capabilities while significantly reducing cloud computing expenses and improving application latency.
Related on Notifire
Related stories
Primary source: Vercel Blog