Protected Content

From Experiment to Trusted Platform

This case study contains detailed work product. Enter the access code to continue.

01 · Executive Overview

Executive Summary

A research prototype with no product owner became the design vision for SnapLogic's entire AI strategy. This is the two-year arc, told the way I'd tell it in a leadership review.

Product / Initiative

SnapGPT, evolved from a standalone assistant into the intelligence layer spanning Agent Creator, MCP Gateway, and platform-wide governance surfaces.

My Role

Director of Product Design, AI/ML. Sole design owner of the AI strategy, with no dedicated researcher on the team for most of this arc.

Business Context

SnapLogic needed a credible AI story in a market where every integration vendor was racing to claim one. Leadership had a technology demo. They didn't have a product.

User Problem

Integration engineers didn't trust or understand the AI capability bolted onto their platform. It lived at the edge of the product. Most didn't know it existed.

Strategic Opportunity

Reposition AI from a bolt-on chatbot to the connective tissue of the platform: the layer that makes automation legible, governable, and worth trusting.

My Contribution

Owned the design vision end-to-end: reframed the product concept, built a research practice from constrained inputs, and negotiated scope across data science, engineering, and product for two years.

Outcomes & Impact

The skill ecosystem became the default entry point to SnapGPT rather than a hidden feature. That reframe earned the design org a seat in the roadmap conversation for Agent Creator, MCP Gateway, and governance, surfaces that didn't exist when this started. The trust patterns built here became the de facto standard the rest of the platform's AI surfaces were designed against.

Visual Language
#0A2647 Chrome navy
#1E4C8C Skill icon blue
#6084A8 Secondary / border
#F5F5F5 Canvas white
Interaction Principles
  • One accent blue, used sparingly, for the nine skill icons and primary actions only, so it stays a signal, not wallpaper.
  • High-contrast white canvas against dark chrome, so the working surface (the pipeline plan) always reads as the primary focus, not the AI panel around it.
  • Circular icon badges for every skill entry point: a consistent shape language that scales as more skills get added without redesigning the grid.
Visual Design Language
Scale

SnapGPT's nine skill icons in Designer are sized as the primary click target, roughly 44 pixels, well above the supporting copy underneath each one. That gap tells a user instantly which layer to scan for options and which to read for detail. The same discipline holds inside Monitor: the forecast's headline, "+112% expected", is the largest text on the panel, while the CPU and memory breakdowns underneath it run at a fraction of that scale.

Visual Hierarchy

Inside Designer, the plan preview renders in full white against the surrounding chat history's dimmer gray, so the one thing that actually requires a decision, approving the plan, never has to compete with the conversational chrome around it. Inside Monitor, the same principle holds in a completely different visual form: the forecast's headline number sits above the manual toggle and the chart, so the insight is always read before the controls used to adjust it.

Balance

When SnapGPT opens inside Monitor, it doesn't take over the screen. The panel is docked to roughly a third of the viewport, deliberately narrower than the node table it sits beside. That asymmetry is the point: SnapGPT is built to augment the surface underneath it, not replace it, and the layout says so before a single word of copy does.

Contrast

Inside the Monitor panel, the only saturated warning color is the single "Over capacity" badge; the "AI forecast" and "Model manually" controls beside it stay in a restrained blue. That restraint is what lets one red badge actually register as urgent instead of competing with a dozen other loud elements for attention.

Gestalt Principles

In Monitor, "AI forecast" and "Model manually" sit inside one shared bordered toggle: common region doing the work of telling a user these two options belong to the same decision, not two unrelated controls. In Designer, the nine skill icons lean on similarity instead: identical circular badges and identical stroke weight tell a user these are peers they can choose between, no matter how different the capabilities behind them are.

02 · Framing the Problem

From a vague technology idea to a usable product experience

Before there was a design problem, there was a definition problem. Nobody had agreed on what SnapGPT actually was.

When I inherited SnapGPT, it had been built by the chief data scientist and an offshore development team, capable, but built to demonstrate what the model could do, not to answer what a user needed. There was no design involvement, no user testing, no product integration. It sat at the edge of the platform, and most users didn't know it existed.

The ambiguity wasn't technical. The model worked. The ambiguity was strategic: was this a chatbot bolted onto an integration tool? A replacement for the visual canvas engineers had built their mental model around for a decade? A research demo that would quietly disappear once the roadmap moved on? Different stakeholders were answering that question differently, and nobody had noticed they disagreed.

The assumption I had to challenge first wasn't about the AI. It was the belief that shipping more chat surface was the same thing as having an AI strategy.

I pushed the team to stop treating "more chat" as the roadmap and start treating trust as the product requirement. That reframe did three things: it moved AI from a feature demo to a design discipline, it gave us a reason to say no to autonomous-by-default patterns that looked impressive but weren't accountable, and it gave leadership a vocabulary (legibility, controllability, accountability) for evaluating AI investment that wasn't just "does the model perform."

The strategic shift, in one line: move from AI as a single chatbot to AI as an intelligence layer that takes whatever shape a given surface actually needs, conversational where the work is exploratory and action based where it's bounded, instead of forcing one unfamiliar paradigm everywhere just to keep the pitch deck simple.

SnapLogic Experience Ecosystem: Creation & Composition, Discovery, Governance, and Observability layers, with SnapGPT's contextual role in each
The full ecosystem, diagrammed. Four layers: creation, discovery, governance, observability. SnapGPT isn't a fifth surface sitting beside them; it's the contextual thread running through all four, which is the actual structural argument behind "intelligence layer, not chatbot."
03 · Decision Stories

Key Design Decisions

Design here wasn't linear. These are the five forks in the road that shaped what SnapGPT, and eventually the platform's whole AI strategy, became.

DECISION 01 One interaction paradigm everywhere vs. one that adapts to the surface
The Fork

The instinct, especially from data science, was to pick one interaction pattern and roll it out identically everywhere SnapGPT appeared: one SnapGPT, one way of talking to it, regardless of surface.

Options Considered
  • A uniform chat-first paradigm, applied identically across Designer, Monitor, and every other surface.
  • A uniform skill/action-menu paradigm, applied identically everywhere, for the smallest possible design surface.
  • An interaction model that adapts to what each surface's tasks actually look like: conversational where the work is exploratory, skill-based where it's bounded and repetitive.
Tradeoffs

A single paradigm is easier to document, market, and build. "One SnapGPT" is a cleaner story than "SnapGPT, but different depending on where you are." But the tasks themselves aren't uniform. Building a pipeline from a blank canvas in Designer is open-ended, and the user often doesn't know the right steps or vocabulary yet, which is exactly the ambiguity a conversation is good at resolving. Reviewing an already-running pipeline in Monitor is the opposite: the tasks are few, well-defined, and repetitive, things like replaying an execution, explaining an error, checking a status. Forcing a conversation onto that adds a step nobody asked for.

Final Direction

Let the interaction model follow the shape of the task, not a platform-wide style guide: conversational, dialogue-driven SnapGPT in Designer; skill- and action-based SnapGPT in Monitor and other bounded surfaces.

Why

The trust principles (a plan the user can read before it runs, a human checkpoint before consequential action) had to hold everywhere. The UI chrome carrying those principles didn't. Once I stopped treating "consistency" as identical pixels and started treating it as identical trust guarantees, the right interaction pattern for each surface stopped being a compromise and became the obvious answer.

SnapGPT open inside Monitor, showing a capacity forecast panel with an AI forecast toggle, a growth slider, and an Update chart button, no chat interface present
SnapGPT in Monitor: capacity forecasting as a skill, not a conversation. There's no prompt box here, just a forecast, a slider to model growth manually, and a button to update the chart. Compare this to Designer, where the same product opens as a chat.
DECISION 02 One generic prompt box vs. a structured skill ecosystem
The Fork

Should SnapGPT be a single freeform prompt box, the generic-chatbot pattern, or a set of named, distinct capabilities users could recognize and reach for?

Options Considered
  • A single freeform prompt box, maximally flexible, minimal design surface.
  • A fixed menu of rigid skills with no freeform input at all.
  • A hybrid: named, discoverable skills, each with freeform language inside its scope.
Tradeoffs

A freeform box feels the most "AI magic" in a demo, but usage data showed people didn't know what to ask and abandoned the empty state. A rigid menu is safe and discoverable, but it caps what power users can do and makes the system feel smaller than the model underneath it actually is.

Final Direction

Nine named skills (documentation, analysis, building, research, help, configuration, and more) each mapped to a real user intent, with natural language inside each one.

Why

The metric I cared about wasn't first-session adoption. It was the moment a user said "I didn't know it could do that" and came back to try something harder. Naming the skills gave people a mental model for what to ask, without capping how they asked it, and that's what turned passive acceptance into expanding engagement.

DECISION 03 Full autonomy vs. tiered, risk-based agent governance
The Fork

As the platform moved toward Agent Creator and genuinely autonomous agents, how much should an agent be allowed to do without a human checking in first?

Options Considered
  • Full autonomy, with an audit log reviewed after the fact.
  • Approval gates on every step, regardless of risk.
  • Tiered autonomy: governance surfaces let admins define what's safe to automate and what requires sign-off, with a full audit trail either way.
Tradeoffs

Enterprise customers, banks and insurers, wanted automation's efficiency, but their compliance functions would never sign off on fully ungoverned autonomous action inside production systems. Approval-gating every step defeats the entire value proposition of agentic automation; nobody adopts an "agent" that pings a human for permission every ten seconds.

Final Direction

Risk-tiered autonomy: routine, reversible actions run unsupervised; irreversible or high-blast-radius actions require a human checkpoint. Audit trail applies to everything, regardless of tier.

Why

This mirrored a principle I'd carried from fraud and healthcare work before AI made it fashionable: humans need to stay accountable for consequential actions. There was a technical reason too: LLM output isn't fully deterministic, so an irreversible action (deleting a pipeline, pushing to production) needed a checkpoint even when 95% of routine actions didn't.

DECISION 04 Pause on research vs. build an alternate research pipeline
The Fork

SnapLogic eliminated the dedicated researcher role partway through this work. Design decisions still needed to be grounded in real usage, not internal opinion.

Options Considered
  • Pause on rigorous research and lean on internal judgment until budget allowed a rehire.
  • Push all research responsibility onto product management.
  • Build a lightweight research pipeline using the customer-facing staff already talking to users daily.
Tradeoffs

Doing nothing is the fastest option, but it risked rebuilding the exact disconnected-from-users problem SnapGPT started as. Waiting for a rehire could mean a year or more without real signal, and the roadmap wasn't going to wait.

Final Direction

Structured conversations with sales engineers and customer advocates, paired with behavioral and adoption data mined from the product itself.

Why

Perfect research access wasn't available, and design couldn't wait for it to appear. Good-enough signal, gathered fast and repeatedly, beat perfect signal that arrived too late to change anything.

DECISION 05 Consolidate ownership vs. embed as the AI design specialist
The Fork

Governance and observability grew from features I'd shaped into full programs with their own dedicated directors. Did I fight to keep them under my ownership, or hand off delivery while staying involved?

Options Considered
  • Consolidate everything under one design leader for a single unified vision.
  • Cede full ownership and step back from those surfaces entirely.
  • Stay embedded as the AI-experience specialist within each program, while another director owned delivery.
Tradeoffs

Consolidating protects a single design vision, but the org was growing two credible programs that each needed dedicated leadership bandwidth I didn't have alongside SnapGPT and Agent Creator. Ceding fully risked governance and observability reverting to engineering-driven UI with no AI-experience point of view at all.

Final Direction

Embedded as the specialist AI designer inside each program, reporting on design quality while another director owned outcomes and delivery.

Why

Scaling impact isn't the same as maximizing headcount under your title. It's making sure the right design thinking reaches the right surfaces, even without direct ownership, while protecting focus on the highest-leverage front I was uniquely positioned to own.

04 · Evidence

Research & Customer Insights

Without a dedicated researcher, the signal came from wherever it actually lived, and I had to build the muscle to hear it.

Three signal sources shaped this work: structured conversations with sales engineers who talked to prospects every week, escalations surfaced by customer advocates when a deal stalled, and usage telemetry showing exactly where people dropped off inside SnapGPT itself.

Signal

Adoption funnel data showed a steep drop-off at the empty "ask anything" state. People opened the assistant, didn't know what to type, and left.

Design Response

Directly informed Decision 02: the shift from a freeform box to a named skill ecosystem that gave people a starting vocabulary.

Signal

Sales engineers reported enterprise deals stalling because prospects' compliance teams couldn't get a straight answer on what the AI was allowed to do unsupervised.

Design Response

Became the business case for Decision 03: the investment in tiered autonomy and a visible audit trail wasn't a nice-to-have, it was a sales blocker.

None of this was a formal study. It was a discipline of asking the same questions consistently, of people who talk to users far more often than any designer gets to, and treating their patterns as data rather than anecdote.

05 · How Thinking Changed

Product Evolution

SnapGPT didn't arrive at its current form in one leap. Each iteration answered a failure of the one before it.

Early SnapGPT exploration board: closed and open sandbar states, a toolbar variant, the teal color ramp, prompt input states, and the feedback modal
Early exploration. Sandbar open/closed states, a toolbar variant we tested and dropped, the color system, and prompt input states across idle, entry, and response generation. The feedback modal in the bottom row is the same mechanism referenced above. It's how signal kept flowing in after the dedicated researcher role was gone.

Iteration 1: contextual embedding, in Designer. Moved the assistant out of an isolated chat surface and into the canvas itself, reading pipeline state so its suggestions were grounded in what the user was actually looking at.

SnapGPT capabilities panel from the original beta release, showing four live capabilities and two marked coming soon
The beta, three years ago. Four live capabilities, two marked plainly as "coming soon." This is the state Iteration 2 below expanded out of, a single, honest capability list, not yet a skill ecosystem.

Iteration 2: the skill ecosystem, still in Designer. As adoption data came in, expanded from a single capability into nine named skills: documentation, analysis, building, research, help, configuration, each progressively disclosed so depth didn't mean overwhelm.

Iteration 3: the platform reframe. Once trust in the assistant itself was established, extended the same design language upward: Agent Creator, MCP Gateway, governance surfaces. The product stopped being "an assistant" and became the orchestration fabric for enterprise AI execution on the platform.

Before

A chat widget buried at the edge of the product. No shared definition of what it was for. Most users didn't know it existed.

After

A skill-aware assistant embedded in the canvas, proposing human-readable plans, backed by governance surfaces that make its autonomy legible and auditable.

06 · Beyond One Feature

Systems Thinking

This work stopped being about one product the moment other teams started asking to reuse the same trust patterns.

Design System

Internal Pattern Library

Documented the trust patterns proven in SnapGPT (plan previews, skill framing, tiered autonomy) as a reusable internal library, so new AI features started from a shared vocabulary instead of reinventing one.

Cross-Program Reuse

Shared Trust Patterns

Audit Trail, Reasoning Disclosure, and HITL Controls patterns built for SnapGPT were adopted by the governance and observability programs, owned by different directors, using the same design language.

Governance

Risk-Tiered Autonomy Model

The autonomy framework from Decision 03 became the default model new agentic features are designed against, not a one-off decision for Agent Creator alone.

Scaling design impact past a single feature meant giving other teams a system they could pick up without me in the room: components with a defined trust rationale, not just a visual style.

07 · The Work

Final Experience

What shipped, and why each choice does more work than it looks like at first glance.

SnapGPT: skill selection and pipeline planning flow
Nine skills, each mapped to a specific user intent
Natural language drives the pipeline, no canvas required
Plan is human-readable before a single snap is placed
My Role Director of Product Design
Problem Emerging AI capabilities had no coherent human-facing layer. Intent, feedback, governance, and accountability were undefined.
Outcome Clearer, more usable product experiences built around intent, automation, system feedback, governance, and accountability.

Why the skill menu works: it's not a navigation convenience. It's the mechanism that turns an intimidating "type anything at the AI" surface into a legible menu of intents, reducing the cognitive load of figuring out what's even possible.

Why the human-readable plan works: it's the single highest-leverage trust mechanism in the whole system. It turns an opaque model decision into something a skeptical integration engineer can read, correct, and approve before anything touches their production pipeline.

Step 1 Read Salesforce records
Step 2 Transform & validate data
Step 3 · Checkpoint Write to production database
Step 4 Confirm write, log result
Step 5 Notify requester
Runs unsupervised: routine, reversible
Requires human sign-off: irreversible or high blast-radius

Risk-tiered autonomy, diagrammed. This is the model behind Decision 03. Most of an agent's pipeline runs unsupervised, but any step with real consequences stops for a human checkpoint before it executes, not after.

09:41:02Agent
Proposed plan: sync Salesforce accounts → BigQuery customer table
Pending review
09:41:18Jordan T. · Admin
Reviewed the plan, approved
Approved
09:41:19Agent
Executed step 1/4: read Salesforce records
Completed
09:41:26Agent
Flagged step 3/4: write to production database (destructive)
Awaiting approval
09:42:03Jordan T. · Admin
Approved the write action
Approved

The audit trail, diagrammed. Every action is attributed to an actor, agent or human, and timestamped. Nothing an agent does disappears into a log line only engineers can read; it's the same record an auditor, a compliance officer, or the admin themselves would need.

Current Scope

SnapGPT · Agent Creator · Agent Visualizer · Prompt Composer · MCP Servers & Agent Tools · Designer Canvas · Admin Manager · Governance Surfaces · Expression Builder · Execution Replay

08 · Looking Back

Reflection

What I Learned

Trust isn't a feeling you design toward. It's a functional state built from legibility, controllability, and accountability. Once I had that vocabulary, every "should the AI do this automatically" debate became answerable.

What I'd Improve

I'd push earlier for a quantitative research budget instead of relying entirely on the alternate pipeline I built. It got us signal fast, but rigorous longitudinal data would have made the case for governance investment sooner.

How This Changed My Thinking

This is where "AI as intelligence layer, not chatbot" stopped being a slogan and became a real design discipline for me, one I now apply to every agentic system I touch.

More Case Studies