← Back to list
Development· 5 min

Work Recap 2026-09

Work Recap 2026-09

AX Challenge

My company held an Augmented Experience (AX) Challenge. The challenge required submission of a document describing how you used AI to improve your work. I joined the challenge as a solo team.

I gave a presentation about my project in company. It won 2th prize award among 20 submissions.

Problem Context

At that time, I was experiencing difficulties with manual testing of software for acquisition.

The number of specified acceptance tests was about 260. 3 people (me, FE, PM) tested independently. It took a lot time.

Every time the software version changed, It needed to be tested again because we couldn't ensure all tests would still pass.

And although we checked that a test case was passed, we couldn't ensure it was actually passed since there were no evidence.

Objective and Approaches

I set 3 objectives.

  • Let an Agent test for me
  • Reduce effort to repeat testing
  • Make tests leave evidence

And I adopted some tech:

  • Playwright for automating web browsing
  • Anthropic Agent SDK for building local Agent
  • Next.js for integrating a visual app and system

Then I started detail design

I designed a self-correction loop: A test case would be handed to Tester Agent. Tester Agent reads it and explores target App by web browsing tool. Tester Agent writes a spec script for testing using playwright and reports a spec script is ready. The system runs it for the Agent. If it fails, the system provides feedback to the Agent and asks to improve it. Repeat these steps until spec script passes.

I added a supervisor loop wrapping the previous loop: When Tester Agent reports any of [APP_DEFECT|CASE_MISS_MATCH|BLOCKED], a Supervisor Agent would analyze the problem. If Supervisor Agent decides the problem is critical and it seem like other case runs could fail, The Agent returns STOP_RUN to the system and reports the problem context to User.

I designed the logic of the system: The system is a dedicated machine launching tests one behalf of the User. It begins all loop logic and orchestrates components inside loops. It provides input (prompts, context, usable tools) for Agents and takes their output (result, structured output, evidence, screenshots, reports), connecting them. It also renders UI and visualizes the loop runs and case run results. The User accesses them easily.

Advance and making it better

I advanced Proof of Concept(PoC). The system worked and had potential for automating the test jobs. But it was not enough to meet actual requirements for staging environments like dev, test.

The works were built in local environment. I isolated each data state before each case running. But web-deployed application could not be isolated directly. So I chose a detour: using random string IDs for every test run to prevent them from affecting each other.

I let Agents leave evidence and logs for their work to optimize performance. It was helpful to find bottlenecks in agent work especially Agent's tool calling log was.

And Tester Agent consumed too many tokens for routine work like login, agreeing terms, setting up a new workspace. It was resource leakage. I provided MCP tools and helper APIs (available to import in spec script) for the Agent. They were designed as functions that set up instead of reading screens and clicking buttons. As a result token consumption decreased by half.

I putted a lot of effort into the design of The functions including MCP tools and helper APIs. It should be matched with general use-cases and Its arguments should be same with User input in the screen. If I let it receive surrogate ids as its arguments, it would have dependency with real implements and would not describe the user's actions. And there were many options to be considered for performance optimization like that which way to implement, what to be open to agents, how to match the behaviors between MCP tool and Helper APIs.

It was not golden path but I isolated detail implements of functions but only opened interface contracts and descriptions for agents. And made their input as same with UI inputs. Those called back-end API instead of using browser to reduce the response time. I let not covered behaviors of those be covered by testing precondition case.

Shifts of Tester Agent's model affected to quality of spec script assertions. I identified a trend that the lower intelligence model wrote more flawed assertions blocks. But using higher one led to over budget so I needed to optimize.

I adopted a Validator agent to review outputs from Tester Agents. I adjusted between cost, quality.

I adopted a Curator agent to improve the loop by oneself. When a case spent many attempts round, the agent improved instruction prompt about screen and locator for the next case loop run.

The next challenges

I've come across Lang Graph, which support for building agent loop and integration agents. And I wanted to used it for developing a Back-end Application.

Comments 0

Be the first to comment.