xAI Introduces Priority Processing, Context Compaction, and Enhanced API Integrations
xAI rolls out several significant API updates, including priority request scheduling, a new context compaction tool, and seamless file sharing between the Files API and Imagine endpoints.
xAI unveils a suite of powerful API updates designed to improve performance and flexibility for developers. Users can now request higher scheduling priority on text, image, and video inference endpoints by setting a specific service tier, ensuring critical tasks receive faster processing while only incurring premium billing when the feature is actively used.
The company also bridges its Files API with the Imagine platform, allowing developers to generate permanent public URLs for stored files and directly reference saved assets in image generation requests without re-uploading. Additionally, a new Context Compaction API enables developers to compress lengthy conversation histories into shorter contexts, which reduces costs, accelerates response times, and sharpens performance in long agent loops.
Further updates include a WebSocket mode for the Responses API that lowers latency during tool-heavy agent workloads, and a Smart Turn feature for the streaming Speech to Text API that uses machine learning to accurately detect when a speaker finishes a thought. Web Search also now supports explicit image searching, returning relevant images directly as Markdown embeds within the response.