
Together AI has introduced a streamlined way for developers to deploy models from Hugging Face using their Dedicated Container Inference (DCI) infrastructure. By leveraging Goose, a CLI agent runner, developers can quickly set up and run models like Netflix's void-model without the usual technical hurdles. This process significantly reduces the time and expertise required to deploy new models, allowing for immediate experimentation and use. The DCI infrastructure provides a private, GPU-backed environment, making it an attractive option for teams looking to quickly integrate new AI models into their workflows.
Read original
© Together AI BlogThunderAgent introduces a novel approach to agentic inference, significantly improving throughput and reducing latency in synthetic data generation. By treating each agent workflow as a program rather than isolated requests, it mitigates KV cache thrashing and balances load across nodes. This results in up to 2.5× higher throughput on single nodes and near-linear scaling on multi-node clusters. ThunderAgent's compatibility with existing inference optimizations makes it a practical choice for enhancing large-scale agentic workloads.
© Together AI BlogTogether AI's partnership with Moonshot AI marks a significant step in making cutting-edge AI models more accessible to developers. By hosting Moonshot's Kimi models, including the 2.8 trillion parameter Kimi K3, Together AI offers developers immediate access to powerful open-weight models. This collaboration allows for seamless integration and post-training capabilities, enabling developers to fine-tune models for specific applications. The partnership promises to deliver high-performance AI solutions with the flexibility and scalability that open models provide, challenging proprietary systems in the market.
© Together AI BlogTogether AI has introduced a sophisticated architecture for model inference that integrates endpoints, deployments, and configurations with capacity-aware traffic splitting. This system allows for seamless rollouts, A/B testing, and zero-downtime updates, making it easier for developers to manage and optimize AI models. By using immutable configurations and a weight-based traffic split, the platform ensures efficient resource allocation and scaling. This development simplifies the deployment process and enhances the reliability of AI applications by ensuring consistent performance and easy rollback options.
© The Verge AIMeta is gearing up for a significant expansion into personal AI agents, aiming to make them accessible and user-friendly for billions of people. CEO Mark Zuckerberg envisions these agents as tools that can assist with various aspects of life, from health to finances, operating seamlessly out of the box. This move differentiates Meta from competitors like Anthropic and OpenAI, which focus more on coding and enterprise solutions. However, Meta faces challenges, including a lack of ecosystem integration and public trust issues. The company plans to reveal more details soon, positioning personal agents as a cornerstone of its future product and revenue strategy.
© The Verge AIPerplexity has extended its Personal Computer tool to Windows, transforming PCs into AI agents capable of managing local files and applications. This expansion follows the Mac version's release and integrates with Microsoft Office 365 and Teams, allowing seamless operation within the Windows environment. The tool is designed for enterprise use, ensuring that AI can assist with tasks traditionally performed on local machines. Available to Max and Enterprise Max users, it emphasizes data privacy by not training on company data and notifying users before executing sensitive actions.
© MIT Technology Review AIIntel is investigating how agentic AI can revolutionize enterprise workflows, moving beyond the capabilities of traditional chatbots. Through extensive experimentation, Intel demonstrates the necessity of focusing on system-wide performance metrics rather than just inference capabilities. Their research underscores the importance of a robust infrastructure that supports scalable systems and precise task orchestration. By emphasizing agent density and task latency, Intel aims to optimize AI performance in business settings. This transition to agentic AI marks a shift from experimental AI to practical, scalable solutions that enhance productivity and governance within enterprises.