Human-in-the-Loop AI

AI Agent Lab

AI Agent Lab

AI Agent Lab

A workflow for generating, evaluating, and improving AI-assisted executive communication.

A workflow for generating, evaluating, and improving AI-assisted executive communication.

A workflow for generating, evaluating, and improving AI-assisted executive communication.

Overview

AI Agent Lab reframes AI-assisted presentation work from a generation shortcut into a structured workflow for communication quality.

The project started with a familiar promise: rough input goes in, a deck comes out. But the more important question was whether the output actually worked — whether it had the right narrative structure, audience posture, communication register, brand fit, and level of editorial judgment.

I designed a modular AI workflow that separates generation, evaluation, and human judgment. The goal was not to replace the designer. It was to make AI-assisted communication easier to structure, review, and improve.

The project produced a documented AI workflow system with separated generation and evaluation layers, reusable narrative skills, sourced brand profiles, evidence-typed findings, and validation runs that show both system strengths and limitations.

System Type

AI Workflow System

System Type

AI Workflow System

Role

I designed the workflow logic, agent roles, handoff architecture, prompt system, evaluation criteria, and human decision points.

Audience

Designers · Communication leads · Executive storytellers · Teams building recurring leadership narratives or evaluating AI-assisted communication quality

Scope

Prototype workflow · Executive Storyline Agent · Independent Review Agent · Shared narrative skills · Sourced brand profiles · Synthetic demo run · Validation runs · Selected evaluation artifacts

Tools / Methods

Claude in VS Code

Microsoft Copilot

PowerPoint

Prompt design

Workflow mapping

Narrative systems design

Register taxonomy

Presentation-quality review

Visual communication design

01 / Problem

AI could generate slides. Generation alone could not tell if they worked.

The problem was not speed. It was communication quality.

The earliest version of the project treated AI as a faster way to create presentations. But the first structured reviews revealed a deeper problem.

Generated decks could look finished while still failing in ways that mattered: the wrong register for the audience, unsafe speaker notes, layout issues that only appeared after rendering, or a story that sounded coherent but missed the communication intent.

That changed the project brief. The goal was no longer to generate better slides. It was to design a workflow that could evaluate whether the communication was actually working.

02 / Separation

Generation and evaluation became separate layers.

Generation and evaluation became separate layers.

Generation and evaluation became separate layers.

A system that creates the output should not be the only system judging it.

Layer 1
Generation

Layer 1
Generation

Creates storyline and presentation structure

Creates storyline and presentation structure

Layer 2
Evaluation

Layer 2
Evaluation

Reviews narrative, register, layout, consistency, and brand fit

Reviews narrative, register, layout, consistency, and brand fit

Layer 3
Human Judgment

Layer 3
Human Judgment

Sets intent, interprets findings,
and makes final decisions

Sets intent, interprets findings, and makes final decisions

Sets intent, interprets findings,
and makes final decisions

The key design decision was to separate production from evaluation.

One layer generates the storyline and presentation structure. A second layer reviews the output without rewriting it. Human judgment stays responsible for strategic intent, editorial decisions, and final readiness.

That separation turned the project from a deck generator into an AI communication workflow.

03 / Workflow

A workflow for moving from rough input to clearer communication.

The system structures ambiguity before it produces artifacts.

The workflow connects raw notes, message prioritization, executive framing, storyline generation, presentation-quality review, brand-aware checks, and human intervention.

Each step has a specific role. Shared skills help structure the message. The generation layer creates the draft artifact. The evaluation layer reviews the output against communication criteria. Human judgment decides what should change.

The result is not a one-click deck. It is a repeatable workflow for turning ambiguous material into clearer executive communication.

Synthetic demo run showing the workflow from structured input to editable output, independent review, and brand-aware QA.

Editable PPTX

Editable PPTX

Independent review report

Independent review report

Brand QA findings

Brand QA findings

04 / VALIDATION

Three validation runs tested the workflow’s range—
not just its best case.

The strongest result showed the system’s potential.
The transfer test revealed its boundary.

The prototype was tested through three runs built around different communication conditions. Rather than selecting a single polished success, the runs examined how the workflow responded to rough input, aligned executive intent, and customer-facing transfer.

01 / STRESS TEST

01 / STRESS TEST

CONDITION

Rough, multi-source input

CONDITION

Rough, multi-source input

SIGNAL

Usable with Fixes

SIGNAL

Usable with Fixes

TEST QUESTION

Could the workflow recover a coherent executive frame from fragmented cross-functional signals?

TEST QUESTION

Could the workflow recover a coherent executive frame from fragmented cross-functional signals?

The workflow recovered a usable narrative arc from fragmented PM, engineering, and sales input. The review distinguished intentional benchmark roughness from structural failure while still surfacing synthesis and prioritization gaps.

The workflow recovered a usable narrative arc from fragmented PM, engineering, and sales input. The review distinguished intentional benchmark roughness from structural failure while still surfacing synthesis and prioritization gaps.

The system could evaluate rough work honestly without treating roughness as total failure.

The system could evaluate rough work honestly without treating roughness as total failure.

02 / ALIGNED EXECUTION

02 / ALIGNED EXECUTION

CONDITION

Clean internal executive brief

CONDITION

Clean internal executive brief

SIGNAL

Review-Confirmed Strong

SIGNAL

Review-Confirmed Strong

TEST QUESTION

Could generation and review converge on a brand-native, register-consistent executive brief?

TEST QUESTION

Could generation and review converge on a brand-native, register-consistent executive brief?

Through documented iteration, the workflow produced the only Strong verdict in the series. Audience intent, narrative structure, communication register, and visual hierarchy remained aligned.

Through documented iteration, the workflow produced the only Strong verdict in the series. Audience intent, narrative structure, communication register, and visual hierarchy remained aligned.

The workflow performed best when audience, intent, and source conditions were explicit.

The workflow performed best when audience, intent, and source conditions were explicit.

03 / BOUNDARY TEST

03 / BOUNDARY TEST

CONDITION

Customer-facing / new brand context

CONDITION

Customer-facing / new brand context

SIGNAL

Intent-Output Mismatch

SIGNAL

Intent-Output Mismatch

TEST QUESTION

Would executive arc logic transfer unchanged to a customer-facing solution story?

TEST QUESTION

Would executive arc logic transfer unchanged to a customer-facing solution story?

The artifact looked polished and used the intended template and brand profile. But the review detected an analytical observer stance rather than a customer-facing persuasive posture.

The artifact looked polished and used the intended template and brand profile. But the review detected an analytical observer stance rather than a customer-facing persuasive posture.

Visual polish and brand alignment did not guarantee communication fit.

Visual polish and brand alignment did not guarantee communication fit.

SYSTEM LEARNING

SYSTEM LEARNING

SYSTEM LEARNING

Quality was not a single visual score.
The runs separated resilience, aligned execution, and transfer limits.

Quality was not a single visual score. The runs separated resilience, aligned execution, and transfer limits.

Quality was not a single visual score.
The runs separated resilience, aligned execution, and transfer limits.

Human judgment remained essential when audience, intent, or communication posture changed.

Human judgment remained essential when audience, intent, or communication posture changed.

Human judgment remained essential when audience, intent, or communication posture changed.

05 / Design Insight

A correct structure can still carry the wrong communication posture.

The strongest boundary test produced a structurally coherent, visually resolved deck—but the language remained analytical when the brief required a customer-facing voice.

The evaluator could identify the mismatch and distinguish what still worked from what needed reframing. Human judgment was still required to decide how the message should feel, what needed reframing, and which parts of the structure should remain intact.

KEY TAKEAWAY

KEY TAKEAWAY

KEY TAKEAWAY

The value of the system was not automatic correctness,
but more precise human intervention.

The value of the system was not automatic correctness,
but more precise human intervention.

The value of the system was not automatic correctness, but more precise human intervention.

06 / Impact

The lab turned one-off generation into a system that can be tested, improved, and extended.

By the end of the prototype, the work had become more than a deck-generation workflow. It became a reusable foundation for narrative generation, evaluation, and iteration.

The system is strongest today in executive narrative work. The next phase is to expand the agent and skill library for customer-facing, product, and recurring communication use cases.

The goal is not broader automation for its own sake. It is a more adaptable system that can match narrative logic, communication posture, and visual structure to the job at hand.

FROM PROTOTYPE TO EXTENSIBLE SYSTEM
The workflow became a reusable foundation for
narrative generation, evaluation, and iteration.

FROM PROTOTYPE TO EXTENSIBLE SYSTEM
One workflow became a reusable foundation for narrative generation, evaluation, and iteration.

FROM PROTOTYPE TO EXTENSIBLE SYSTEM
The workflow became a reusable foundation for
narrative generation, evaluation, and iteration.

CURRENT FOUNDATION

CURRENT FOUNDATION

01 / NARRATIVE SYSTEM

01 / NARRATIVE SYSTEM

Reusable narrative logic

Reusable narrative logic

Message prioritization

Message prioritization

Executive framing

Executive framing

Narrative architecture

Narrative architecture

Structured handoffs

Structured handoffs

02 / SPECIALIZED AGENTS

02 / SPECIALIZED AGENTS

Separated responsibilities

Separated responsibilities

Executive Storyline Agent

Executive Storyline Agent

Independent Review Agent

Independent Review Agent

Human decision points

Human decision points

03 / QUALITY EVIDENCE

03 / QUALITY EVIDENCE

Traceable evaluation

Traceable evaluation

Evidence-typed findings

Evidence-typed findings

Versioned validation runs

Versioned validation runs

Sourced brand context

Sourced brand context

Documented strengths and limits

Documented strengths and limits

Quality became something the workflow could examine—not simply assume.

Quality became something the workflow could examine—not simply assume.

Quality became something the workflow could examine—not simply assume.

NEXT BUILD

NEXT BUILD

NARRATIVE EXPANSION

NARRATIVE EXPANSION

Expand what the system can communicate

Expand what the system can communicate

Customer-facing solution stories

Customer-facing solution stories

Product and transformation narratives

Product and transformation narratives

Recurring leadership communications

Recurring leadership communications

Use-case-specific agents and register skills

Use-case-specific agents and register skills

LAYOUT EXPANSION

LAYOUT EXPANSION

Expand how the system can structure information

Expand how the system can structure information

Expand how the system can
structure information

Broader layout archetype library

Broader layout archetype library

Content-to-layout matching

Content-to-layout matching

Reusable layout skills

Reusable layout skills

Layout-fit review criteria

Layout-fit review criteria

The next phase expands both what the system can say and how it can show it.

The next phase expands both what the system can say and how it can show it.

The next phase expands both what the system can say and how it can show it.

Evidence Boundary


What this case study demonstrates—
what it does not claim

This case study documents a working prototype, reusable system components, and a series of structured validation runs. The evidence demonstrates how the workflow generates, evaluates, and improves communication across different conditions.

It does not claim production deployment, organization-wide adoption, or measured business impact.

CONFIDENTIALITY NOTE

Selected examples use synthetic or portfolio-safe content where necessary. Confidential source material, customer information, and internal implementation details have been omitted.