SIGNAL
Tracking the global AI frontier — labs · research · agents · policy
Frontier Signal

Agents

The agentic stack, tracked daily: MCP and A2A protocols, LangChain/LangGraph, Claude Code, OpenAI Agents SDK, agent frameworks, benchmarks and orchestration patterns.

The agentic stack, tracked daily: MCP and A2A protocols, LangChain/LangGraph, Claude Code, OpenAI Agents SDK, agent frameworks, benchmarks and orchestration patterns.

Latest in Agents

30 stories
Shared infrastructure, isolated tenants: Pool model multi-tenancy with Amazon Bedrock AgentCore
Practice

Shared infrastructure, isolated tenants: Pool model multi-tenancy with Amazon Bedrock AgentCore

In this post, you will learn patterns for implementing production-ready multi-tenant systems using Amazon Bedrock AgentCore. You will see these patterns demonst…

How Businesses Are Building Specialized AI They Can Trust
Labs

How Businesses Are Building Specialized AI They Can Trust

Companies are asking how to build specialized AI that fits with the way their workflows actually run. The first wave of enterprise AI was about access. Companie…

GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
Research

GroundEval: A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation

Before letting an agent operate over real context, can you prove it used the right evidence? GroundEval turns that question into a deterministic test of what th…

RaMem: Contextual Reinstatement for Long-term Agentic Memory
Research

RaMem: Contextual Reinstatement for Long-term Agentic Memory

Long-term memory has become increasingly important for LLM agents that operate across extended interactions and evolving task contexts. Recent memory systems ha…

AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions
Research

AI Scientists as Engines of Discovery: A Case for Development within Reformed Institutions

Agentic artificial intelligence (AI) systems are beginning to assist, accelerate, and partially automate scientific discovery, performing tasks that span litera…

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO
Research

IPO Finance Agent: Evaluation of LLM Financial Analysts beyond Finance Agent v2, with Automated Rubric Generation -- the Case of the SpaceX (SPCX) IPO

Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language models on financi…

Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning
Research

Self-Evolution for Multi-Turn Tool-Calling Agents via Divergence-Point Preference Learning

Multi-turn tool-using agents must coordinate long-horizon tool sequences while tracking dialogue state and policy constraints. Existing approaches often separat…

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation
Research

RS-Gen: A Multi-Stage Agentic Framework for Reasoning and Search-Augmented Image Generation

Recent years have witnessed remarkable progress in image generation and editing, particularly regarding instruction following and visual fidelity. However, when…

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction
Research

AOHP: An Open-Source OS-Level Agent Harness for Personalized, Efficient and Secure Interaction

AI agents are driving a new software paradigm, with the ability to autonomously call tools, extract information, manage memory, and complete tasks that span app…

Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference
Research

Concordia: JIT-Compiled Persistent-Kernel Checkpointing for Fault-Tolerant LLM Inference

Long-running LLM agents keep valuable state resident on GPUs: KV caches, request schedulers, communication state, and sometimes online adapters. Losing this sta…

NVIDIA Brings Trusted, 24/7 AI Agents to Telecom Operations
Labs

NVIDIA Brings Trusted, 24/7 AI Agents to Telecom Operations

Telecom operators have seen remarkable returns from using generative AI to automate network management, customer care and back-office operations. Most of that i…

Google makes Interactions API the default interface for Gemini models and agents
Briefing

Google makes Interactions API the default interface for Gemini models and agents

Google Deepmind has made the Interactions API the default interface for Gemini models and agents. It replaces the old generateContent API and uses a simplified …

Building pay-per-intelligence for AI agents: How Ampersend uses Amazon Bedrock AgentCore Payments
Practice

Building pay-per-intelligence for AI agents: How Ampersend uses Amazon Bedrock AgentCore Payments

In this post, you will learn how Ampersend built a pay-per-intelligence routing layer on top of Amazon Bedrock AgentCore Payments. AI agents autonomously route …

New features and Claude as agent provider preview in JetBrains IDEs
Create

New features and Claude as agent provider preview in JetBrains IDEs

This update adds support for organization and enterprise agents from GitHub, lets you queue and steer messages in Copilot CLI sessions, introduces a new agent d…

GLM-5.2 is the step change for open agents
Research

GLM-5.2 is the step change for open agents

A capability threshold I've been carefully monitoring.…

CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?
Research

CODA-BENCH: Can Code Agents Handle Data-Intensive Tasks?

Advanced agents are increasingly demonstrating the potential to operate as autonomous engineers, creating a growing demand for evaluation benchmarks that captur…

Repeated Bilateral Trade: The Quest for Fairness
Research

Repeated Bilateral Trade: The Quest for Fairness

We study repeated bilateral trade from a fairness perspective. At each round, a fresh seller-buyer pair arrives, and the platform posts a price before observing…

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment
Research

Defending against Adaptive Prompt Injection Attacks via Reasoning-enabled Task Alignment

Indirect prompt injection attacks hijack LLM-based agents by embedding malicious instructions in third-party data that the agent retrieves during task execution…

LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis
Research

LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis

Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and hi…

Can In-Context Learning Support Intrinsic Curiosity?
Research

Can In-Context Learning Support Intrinsic Curiosity?

Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized da…

LooseControlVideo: Directorial Video Control using Spatial Blocking
Research

LooseControlVideo: Directorial Video Control using Spatial Blocking

Precise 3D spatial orchestration in text-to-video generation remains a significant challenge, particularly for multi-object scenes where semantic layout and tem…

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning
Research

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agent…

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View
Research

See-and-Reach: Precise Vision-Language Navigation for UAVs within the Field of View

UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target appr…

Eco Wave Power Turns Waves Into Watts With NVIDIA AI Infrastructure and Digital Twins
Labs

Eco Wave Power Turns Waves Into Watts With NVIDIA AI Infrastructure and Digital Twins

The next era of AI will not be defined by compute alone. Its growth will be determined by energy. As accelerated computing scales across AI factories, agentic A…

NVIDIA Vera CPU Opens the Way for Agentic Scientific AI at Los Alamos National Laboratory
Labs

NVIDIA Vera CPU Opens the Way for Agentic Scientific AI at Los Alamos National Laboratory

Mission, Vision and Veritas — new Los Alamos National Laboratory (LANL) supercomputers to be built with HPE and NVIDIA — are tapping NVIDIA Vera CPUs to acceler…

From campaigns to continuous growth: AI capabilities shaping marketing
Practice

From campaigns to continuous growth: AI capabilities shaping marketing

Capabilities built around insights, creativity, personalization, agentic commerce, and orchestration form the five core pillars of the future of marketing.…

Sakana Fugu: One Model to Command Them All
Labs

Sakana Fugu: One Model to Command Them All

Sakana AI Releases ‘Fugu Ultra’ to Match Frontier Performance via Autonomous Model Orchestration.Our Fugu Ultra model stands shoulder-to-shoulder with leading m…

AWS says AI agents lack business context and security, launches two services to patch the gaps
Briefing

AWS says AI agents lack business context and security, launches two services to patch the gaps

At its summit in New York, AWS unveiled two new services. Continuum automatically detects, prioritizes, and fixes code vulnerabilities. Context builds a knowled…

Data2Story turns a CSV file into a verified interactive news article using seven AI agents
Briefing

Data2Story turns a CSV file into a verified interactive news article using seven AI agents

Seven AI agents work together like a newsroom. The "Data Journalist Agent" from Oxford and Stanford turns a CSV file into a finished interactive article with gr…

A Human-Augmenting Agentic Workflow for Causal Inference
Practice

A Human-Augmenting Agentic Workflow for Causal Inference