Skip to content
Back to work

AI & Automation · Healthcare

Autonomous AI Log Monitoring & Observability Platform

For a healthcare media and clinician engagement platform, a five-agent system that reads a production error, writes the fix and opens a reviewed pull request, with an engineer still deciding what ships.

Role
AI Product Manager & Solution Architect
Built on
Google Cloud, Agent Development Kit, Claude via Vertex AI
Industry
Healthcare
Project type
AI & Automation
Primary service
AI & Automation
Scope
Product ownership, solution architecture, agent workflow design, observability dashboard, delivery.

Repository not shown as per company policy

  • 0

    specialized agents in one workflow, each with a single job

  • ±0min

    of logs read around every error during triage

  • 0k

    token cap per incident, enforced live by a circuit breaker

  • 0

    trace ID follows an incident from alert to pull request

01Challenge

The slow part of an outage is the middle.

A modern microservice estate produces thousands of log lines a minute. When something breaks, an on-call engineer has to notice the alert, find the right logs, reconstruct the failing request, locate the responsible file and commit, write and test a fix, get it reviewed and ship it. The middle steps are mechanical, and they are where most of the hours go.

  • Context is scattered

    Logs, source code and deployment state live in three different systems. Someone has to stitch a trace ID to a file to a commit by hand.

  • Tools only report

    Traditional dashboards show that something broke, not why, and they never propose a fix.

  • The same bugs come back

    Null checks, type mismatches and timeouts get re-diagnosed from scratch by whoever is on call that week.

02Approach

Instrument once. Let agents do the mechanical work. Keep people in charge of the decision.

  • Make every incident traceable

    A standard tracing library across all services means an agent can always start from one ID and find everything related to a failure.

  • Split the work across specialists

    Triage, code retrieval, fix writing, review and Git operations are separate agents with strict, typed hand-offs, each easy to test and improve independently.

  • End at a pull request

    The output is a reviewed proposal in the tool engineers already use. The system prepares work; a person merges it.

03Outcome

An engineer's first draft, written by the platform.

When a production service throws an error, the platform does more than raise an alert. It reads the logs, finds the responsible code, writes a fix, has a second agent check it, and opens a pull request for an engineer to review.

  • Shared tracing

    One library stamps every log line in every service with a trace ID, so any incident can be followed end to end.

  • Remediation swarm

    Five specialized agents move from alert to pull request, with safety checks between each step.

  • Observability dashboard

    A single screen shows what the system did, why, what it changed and what it cost.

What changed

  • Engineers receive a reviewed pull request with a root-cause explanation, rather than only an alert.
  • Every automated change is checked by a separate review agent before it reaches source control, so the author of a change is never its only reviewer.
  • Token spend is capped per incident and reported per agent, so an autonomous run cannot produce an open-ended bill.

How it works

Watch an incident travel through the system.

Services log to centralized logging. An error alert reaches a webhook that starts the agent swarm, which reads code from and opens pull requests in source control. A dashboard reads the resulting audit trail.

runs in backgroundreads the audit trailEvery stepis loggedreads code, opens PRsBusiness microservicesNode and Python servicesusing one tracing libraryCloud LoggingStructured JSON, every linestamped with a trace IDERROR alert + Pub/SubLog-based alert pushesthe event onwardWebhook receiverSkips events withno trace IDObservability dashboardNext.js, behind Google IAPTimeline, diff and costEngineers and on-callReview, approveand mergeBitbucket CloudSource code andpull requestsAutonomous remediation swarmGoogle ADK on Cloud Run, Claude via Vertex AI1 Log AnalystFinds service, file and severity2 EnvironmentFetches code at the running commit3 CoderWrites the corrected file4 ReviewerIndependent security and syntax check5 GitOpsBranch, commit and pull requestToken watchdogCircuit breaker at 30,000 tokens

Every step writes to the audit trail, so the dashboard can reconstruct any incident after the fact without a live session.

From a stack trace to a pull request.

01 / 08

Illustrative walkthrough of a typical incident.

Inside the system

What it is made of, and what keeps it safe.

Specialists with one job each.

Each agent reads defined inputs and must return a strict, typed result. The next agent only runs if the previous one succeeded.

  • 1

    Log Analyst

    Reads five minutes of logs either side of the error, names the service, file and commit, and rates business criticality as low, medium or high.

    GuardrailA fixed rubric maps error types to criticality. Authentication failures, payment services and database disconnects rank high.

  • 2

    Environment Agent

    Retrieves the exact source file from version control at the commit that was actually running in production.

    GuardrailIf the logs do not state a commit, it falls back to the currently deployed one rather than guessing.

  • 3

    Coder Agent

    Writes the corrected file and a short explanation of what changed and why.

    GuardrailIt must keep the file architecture intact and add a comment above the fix.

  • 4

    Review Agent

    Gives an independent second opinion: no new security issues, valid syntax, and the original bug actually resolved.

    GuardrailA separate persona from the Coder, so the author of a change is never its only reviewer.

  • 5

    GitOps Agent

    Creates a branch, commits the fix and opens a pull request that explains the root cause.

    GuardrailIt may only use the repository and file the orchestrator supplies. It never invents names.

  • ↻

    Orchestrator

    Wires the five agents into one workflow, passes state from each to the next, and stops the run the moment a stage fails to report success.

    Why it mattersA weak diagnosis can never slide through into code being written.

Built with

What it runs on.

AI orchestration
Agent Development KitTyped output schemas per agent
Model
ClaudeVertex AI
Backend
PythonFastAPIPydantic
Tracing and logging
OpenTelemetryW3C Trace ContextTypeScript SDKPython SDK
Dashboard
Next.jsReactRedux ToolkitRecharts
Eventing
Cloud Logging alertsPub/Sub push subscription
Security
Identity-Aware ProxySecret Manager
Delivery
Cloud BuildArtifact RegistryCloud RunImages tagged by commit

Roadmap

What the platform is designed to grow into.

Prioritized the way a product owner would: control first, then reach, then polish.

Human in the loop

  • Alerts to chat, email or the issue tracker when the system is stuck or a fix is rejected
  • A confidence signal, so uncertain diagnoses go to a person before any code is written
  • A one-click approval screen for pending pull requests

Safety rails

  • A maintained service-to-repository map that stops and asks when unsure
  • Duplicate-alert protection, with branches named after the trace ID
  • A risk policy where high-criticality services always require human review

Traceability and experience

  • Trace IDs created at the click in the dashboard, so a manual resume is auditable end to end
  • A cleanup view for stale branches and pull requests
  • Confidence flags surfaced directly in the incident detail view

A production error in. A reviewed pull request out. People in control of what ships.

More work

Related projects.

AI & AutomationHealthcare

Drug Competitor Identification

A brand-intelligence tool that asks a language model who a drug competes with, then checks the answer against regulatory reference data before anyone is asked to trust it.

AI & AutomationHealthcare

AI Agents Platform

Seven agents that turn one upload into a recorded, print-ready batch of personalized posters, with a reviewer approving anything that carries commercial risk.

AI & AutomationTechnology & SaaS

Customer Churn Prediction Model

For a subscription software business, a churn model that flags drifting accounts early enough for the customer success team to act, with the reasons behind every score.

Discuss a similar project.

If something here is close to what you need, tell us about your situation and we will walk you through how we would approach it.

Intelligence → Innovation → Automation → Growth