Agents
The agentic stack, tracked daily: MCP and A2A protocols, LangChain/LangGraph, Claude Code, OpenAI Agents SDK, agent frameworks, benchmarks and orchestration patterns.
The agentic stack, tracked daily: MCP and A2A protocols, LangChain/LangGraph, Claude Code, OpenAI Agents SDK, agent frameworks, benchmarks and orchestration patterns.
Latest in Agents
30 stories
Shared infrastructure, isolated tenants: Pool model multi-tenancy with Amazon Bedrock AgentCore
In this post, you will learn patterns for implementing production-ready multi-tenant systems using Amazon Bedrock AgentCore. You will see these patterns demonst…

How Businesses Are Building Specialized AI They Can Trust
Companies are asking how to build specialized AI that fits with the way their workflows actually run. The first wave of enterprise AI was about access. Companie…
GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what th…
RaMem: Contextual Reinstatement for Long-term Agentic Memory
Long-term memory has become increasingly important for LLM agents that operate across extended interactions and evolving task contexts. Recent memory systems ha…
AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions
Agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span litera…
IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO
Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language models on financi…
Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning
Multi-turn tool-using agents must coordinate long-horizon tool sequences while tracking dialogue state and policy constraints. Existing approaches often separat…
RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when…
AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span app…
Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference
Long-running LLM agents keep valuable state resident on GPUs: KV caches, request schedulers, communication state, and sometimes online adapters. Losing this sta…

NVIDIA Brings Trusted, 24/7 AI Agents to Telecom Operations
Telecom operators have seen remarkable returns from using generative AI to automate network management, customer care and back-office operations. Most of that i…

Google makes Interactions API the default interface for Gemini models and agents
Google Deepmind has made the Interactions API the default interface for Gemini models and agents. It replaces the old generateContent API and uses a simplified …

Building pay-per-intelligence for AI agents: How Ampersend uses Amazon Bedrock AgentCore Payments
In this post, you will learn how Ampersend built a pay-per-intelligence routing layer on top of Amazon Bedrock AgentCore Payments. AI agents autonomously route …
New features and Claude as agent provider preview in JetBrains IDEs
This update adds support for organization and enterprise agents from GitHub, lets you queue and steer messages in Copilot CLI sessions, introduces a new agent d…

GLM-5.2 is the step change for open agents
A capability threshold I've been carefully monitoring.…
CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that captur…
Repeated Bilateral Trade: The Quest for Fairness
We study repeated bilateral trade from a fairness perspective. At each round, a fresh seller-buyer pair arrives, and the platform posts a price before observing…
Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution…
LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis
Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and hi…
Can In-Context Learning Support Intrinsic Curiosity?
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized da…
LooseControlVideo: Directorial Video Control using Spatial Blocking
Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and tem…
Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agent…
See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View
UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target appr…

Eco Wave Power Turns Waves Into Watts With NVIDIA AI Infrastructure and Digital Twins
The next era of AI will not be defined by compute alone. Its growth will be determined by energy. As accelerated computing scales across AI factories, agentic A…

NVIDIA Vera CPU Opens the Way for Agentic Scientific AI at Los Alamos National Laboratory
Mission, Vision and Veritas — new Los Alamos National Laboratory (LANL) supercomputers to be built with HPE and NVIDIA — are tapping NVIDIA Vera CPUs to acceler…
From campaigns to continuous growth: AI capabilities shaping marketing
Capabilities built around insights, creativity, personalization, agentic commerce, and orchestration form the five core pillars of the future of marketing.…
Sakana Fugu: One Model to Command Them All
Sakana AI Releases ‘Fugu Ultra’ to Match Frontier Performance via Autonomous Model Orchestration.Our Fugu Ultra model stands shoulder-to-shoulder with leading m…

AWS says AI agents lack business context and security, launches two services to patch the gaps
At its summit in New York, AWS unveiled two new services. Continuum automatically detects, prioritizes, and fixes code vulnerabilities. Context builds a knowled…

Data2Story turns a CSV file into a verified interactive news article using seven AI agents
Seven AI agents work together like a newsroom. The "Data Journalist Agent" from Oxford and Stanford turns a CSV file into a finished interactive article with gr…
