Anyone who has actually deployed an AI agent knows that operating one is harder than building one. No matter how carefully you test beforehand, once real traffic starts flowing in, cracks appear in places you didn't expect. You end up facing an agent that retrieves the wrong context, skips a necessary step, or builds an answer on reasoning that looks plausible but is actually wrong. How do you detect this, how do you fix it, and how do you turn it into a repeatable improvement process? This is the question AI teams around the world are wrestling with right now, and on the ground this year, we saw the industry's answer converging on a single direction.
On June 4th, CloudNetworks' Strategic Business Division and AI Innovation Team attended Arize Observe 2026 in person at Shack15 in San Francisco's Ferry Building. As an official Arize AI partner, we listened especially closely, knowing that the questions this event addresses are exactly the ones our customers are facing. After spending the full day there, the conclusion we reached was surprisingly simple: agents are heading toward an era of self-improvement, and the human role is to measure and safeguard whether that direction is the right one. This article is a record of the day that led us to that conclusion.
Β
ποΈ What Is Arize Observe 2026

Arize Observe is Arize AI's annual conference on AI observability and agent evaluation, held every year in San Francisco. This year marked its fifth edition, held under the name "The AI Agent Evals Conference." 2026 sponsors included AWS, Microsoft, Swift Ventures, CrewAI, Quality Kiosk, and Band, with technology leaders from Anthropic, Cursor, OpenAI, PromptQL, WorkOS, Factory, and Daytona taking the stage as speakers.

The venue, Shack15, sits inside San Francisco's Ferry Building, and the atmosphere was distinctive from the moment you walked in. Attendees naturally opened their laptops and kept conversations going under the arched glass ceiling, and off to one side of the lobby stood a large cube installation engraved with the phrase "THE AGENT FEEDBACK LOOP." Circular stickers reading "Trace it. Evaluate it. Fix it. | Arize AI" were laid across the floor. It was a way of expressing what the event was about through the space itself.
With three parallel tracks running simultaneously β The Arena, Cerebral Valley, and SoMa β planning a route in advance mattered, and doing so turned out to be decisive in making the most of the day. It also helped to review last year's Observe sessions the night before, mapping out what threads might continue this year.
Β
π€ Keynote: The Direction Toward Self-Improving Agents

The keynote opened with Arize AI co-founder and CEO Jason Lopatecki and co-founder and CPO Aparna Dhinakaran, followed by SallyAnn from Product and Roger from Open Source/Phoenix. There was one central declaration: "2025 was the year of the agent."
AI coding agents represented by Claude Code, Cursor, and Codex β a category Arize calls "Harness" β became genuinely practical work tools for the first time, marking the moment agents stepped out of the lab and into real production environments. It was striking to hear that a term that didn't even exist a year ago has now become the most important product category.
The vision the keynote laid out was the self-improving agent. Today, humans still review traces directly and evaluate and improve them manually, but through Arize's automation β the Alyx copilot, automated evaluation, and Skills β the loop that humans have to run by hand is gradually shrinking. The ultimate goal is a structure where a fleet of agents in the cloud handles issue discovery, fixes, and review on its own. In this picture, the human role is redefined as confirming that agents are doing the right thing and controlling the system so it keeps getting better.
One thing became clear here: more automation means a split between what humans let go of and what they hold onto more tightly. What should be let go of is "execution." What must be held onto is "judgment." Nearly every session that followed that day ultimately converged on this same axis.
π New Arize AX Features Unveiled at Observe 2026

In support of the keynote's vision, Arize AI also unveiled a major round of feature updates to Arize AX at the event. The core direction is to unify the entire agent feedback loop β from failure detection through root-cause investigation, fix verification, and continuous improvement β into a single engineering workflow. Following the announced features, it becomes clear they all point toward the same thing: handing off, one step at a time, execution work that humans used to do themselves.
- Signal continuously reviews production traces, automatically detects recurring failure patterns, and produces investigation reports containing root-cause analysis and recommended actions. Rather than digging through traces after a problem has already surfaced, it's structured to surface issues before they pile up β effectively automating the work of finding issues.
- Agent Orchestration runs managed agents with repository access, allowing teams to delegate work such as failure investigation, code analysis, eval generation, fix proposals, and security issue review. It supports Claude Code Managed Agents, Vercel, and Daytona, and changes an agent proposes are applied only after an engineer reviews and approves them. The agent carries the work from issue discovery through to a proposed fix and a generated PR, while final judgment and approval remain in human hands.
- Harness-as-a-Judge doesn't rely solely on predefined evaluation criteria; it automatically generates evaluation signals tuned to failure patterns newly emerging in production. This means agents are now starting to assist even with evaluation β the area that, until now, has required the most human involvement. Even when unexpected failures appear, the evaluation framework can keep pace.
- Full-Agent Experimentation goes beyond prompt-level testing to compare the behavior of an entire agent system β tool-use patterns, retrieval quality, latency, traces, and evaluation results β at the level of full runs, making it possible to confirm through actual execution whether a change that fixed one thing broke something else.
- Voice Agent support extends the observability workflow that has been applied to text agents into voice conversation systems, allowing teams to review audio sessions, transcripts, and multimodal traces together, and to play back or directly evaluate voice conversations.
Β
π£οΈ Key Sessions from the Floor
Because three tracks ran at the same time, we planned our route from the start.

Anthropic's Marius Buleandra presented on building trustworthy agents on top of frontier models β the message that human calibration remains irreplaceable even as automation accelerates stood out. As models get smarter, calibrating whether their judgment is heading in the right direction becomes, if anything, even more important. OpenAI's Stuart Sy covered an approach for turning noisy customer feedback into structured trust signals, and Cursor's John Gilhuly shared how they run an agentic coding SDLC and operate remote agents. Uber's Aayush Agrawal, in a talk titled "The Hardest Part of Evals Isn't the Tooling," addressed the real-world difficulty of embedding an evaluation framework into an organization. Salesforce introduced a methodology for continuously validating the behavior of large-scale multi-agent systems using judge, critic, and simulator agents, and CVS Health presented an operational framework for turning demos into sustainable production ROI.

We made a point of catching the financial-services and regulated-industry sessions separately. Wells Fargo shared its experience building governed agentic AI within a regulated environment. BlackRock's Abhigya Jain gave a particularly memorable Responsible AI session, with the core message that "Responsible AI is a product, not a feature" β the view that safeguards shouldn't be an option bolted on after the fact, but must be designed in as part of the product from the start. The talk walked through a guardrail architecture covering PII leakage, prompt injection, hallucination, unauthorized investment advice, and bias mitigation, and emphasized that this framework is applied consistently across model development, AI tools, chat, and the platform as a whole. For any organization that has to operate AI within a regulated environment, the content was concrete enough to use as a direct reference.
Β
π€ Details That Felt Distinctly Arize
Walking around the venue, what the company builds came through clearly even in the side programming.

The biggest crowd-pleaser was a caricature booth where AI drew attendees' faces in real time β a robotic arm literally held a pen, completed the sketch, and printed it out alongside the Arize logo. It was popular enough that the line never let up. At an AI observability conference, watching AI draw a picture in person felt pretty symbolic in itself.

A booth called "Human-in-the-loop" was also memorable β a space where attendees could pick out AI-themed custom patches and attach them to a sweater themselves. It was a fun way of translating the event's message β that a human adds the final touch β into a hands-on merch experience. Rather than just hanging the concept on a banner, making attendees experience it hands-on felt distinctly Arize.
Β
βοΈ A Korean Company's Case Study on the Global Stage: LG U+'s AICC Presentation

The session we listened to most closely at this event was LG U+'s presentation, titled "from Callbot to AI Agent: How LG U+ Reinvented Customer Service for 30M Subscribers" β a session on their journey from callbot-based operations to AI agents.
The architecture shared in the talk had AI transcribing calls in real time, retrieving answers from a knowledge base and surfacing them on the agent's screen, automating call quality monitoring, and routing to domain-specific expert models through a multi-agent structure. It was also emphasized that general-purpose benchmarks don't guarantee quality within a company's own domain, and that a customer- and domain-specific evaluation framework is essential.
The message running through the talk was clear: "If you can measure, you can develop." You need a measurable evaluation framework in place first in order to know which direction improvement should take β and only then can you start simple and scale from there. It's a message that applies to any organization considering AI adoption. Establishing measurement criteria ultimately means making it clear, even as automation accelerates, what humans will look at to judge whether something has actually improved.
The fact that a Korean company presented its own AI agent operations case study on Arize AI's global conference stage was itself a sign that the level of enterprise AI operations in Korea already holds a seat at the global conversation.
Β
π The Outdoor After-Party by the Ferry Building

After all the sessions wrapped up, Observe After Hours continued in the outdoor space of a nearby hotel. With the Bay Bridge as a backdrop, attendees gathered in small groups to talk β and that time may well have held the most candid conversations of the whole day. The threads from the day's sessions carried on in a more relaxed atmosphere, and it wasn't unusual to see people who'd just met making plans for a follow-up conversation. It was a reminder that, on a conference trip, happy hour isn't optional.
Β
π§ What Arize Observe 2026 Left Us With
One sentence captures what ran through Observe 2026:
"Agents are heading toward an era of self-improvement. The human role is to measure and safeguard whether that direction is the right one."
The keynote's Signal and Agent Orchestration automating everything from issue discovery to fix and PR generation, and Harness-as-a-Judge handing off even evaluation to agents β all of it pointed to the same place. As agents come to handle more and more work on their own, what humans need to let go of is "execution," and what they need to hold onto is "judgment." BlackRock saying "Responsible AI is a product, not a feature," and Anthropic emphasizing that "human calibration is irreplaceable" β in the end, these were saying the same thing. The faster automation moves, the more important it becomes to measure and safeguard whether that automation is heading in the right direction.
This same question is being raised in domestic customer environments as well: how to evaluate, how to quickly find where things are failing, and where to place human judgment while still turning the improvement loop into a repeatable process. As an official Arize AI partner, CloudNetworks will support customers in translating the direction and new capabilities confirmed at this event into an adoption plan suited to their own environment. If you're running LLM-based services, RAG pipelines, or AI agents in production and need to build out quality management and evaluation frameworks, please feel free to reach out.
Β
βΆ Learn more about Arize AI