The AI Engineering Iceberg: What It Really Takes to Ship AI to Production

Anyone can build an AI demo. Connect an API, write a clever prompt, wrap it in a chat window — by Friday you have something that impresses a room. I’ve watched teams do it in a weekend.

Then they try to put it in front of real users, and the iceberg shows up.

Below the waterline of every AI demo sits roughly 90% of the actual engineering: the data pipelines, the retrieval tuning, the evals, the guardrails, the cost controls, the incident response. None of it is visible in the demo. All of it decides whether the thing survives contact with production.

There’s a popular visual making the rounds — Jean Lee’s AI Engineering Iceberg — that captures this perfectly: a small tip above water (prompts, API calls, chatbots) and a massive submerged body of everything else. What follows is my version of what’s down there, drawn from years of shipping data and AI systems at enterprise scale.

The visible tip

Prompt engineering. Calling LLM APIs. Building chatbots.

This is real work, and it’s where everyone starts. It’s also maybe 10% of a production system. The tip is the part you can learn in a weekend and demo on a Friday. Everything below takes the other 90% of the calendar — and nearly 100% of the budget that matters.

The tip is also where most AI strategies stop. “We need an LLM strategy” usually means “we built a chatbot.” The organizations that get past this stage are the ones that ask what happens after the demo.

Below the waterline: data

Every production AI system I’ve seen fail failed on data first.

Data pipelines — ingestion, cleaning, chunking, enrichment. The demo reads three PDFs. Production ingests thousands of documents in five formats, half of them scanned, with tables that break every parser you try.

Retrieval engineering — hybrid search, reranking, metadata filtering. “RAG basics” is chatting with documents. Retrieval engineering is making the right document win, every time, when the corpus is messy and the questions are adversarial. This is where most RAG pilots quietly die: the retrieval looks fine on ten test queries and falls apart on ten thousand real ones.

If your data foundation was built for dashboards, your agents are standing on a surface designed for human eyes. I wrote about this at length in the AI-ready data field guide — the short version is that data good enough for humans is not automatically good enough for machines.

Below the waterline: quality

Model evaluation — quality, accuracy, relevance, task success. The demo has no evals; it has a founder nodding at outputs. Production needs golden datasets, regression suites, and a definition of “good” that survives contact with edge cases.

Observability — tracing prompts, tools, tokens, failures. When a user reports a wrong answer on Tuesday, you need to replay exactly what the system saw, retrieved, and decided. Without traces, debugging an AI system is archaeology.

If you can’t measure it, you can’t ship it — you’re just hoping.

Below the waterline: systems

Agent orchestration — tools, workflows, memory, multi-agent coordination. The demo calls one model. Production orchestrates dozens of moving parts that fail independently and need to fail gracefully.

Tool calling — schemas, permissions, retries, validation. Giving an LLM a tool is the tip. Making that tool call reliable, permissioned, idempotent, and auditable is the iceberg.

Structured outputs — reliable JSON, parsing, schema enforcement. Downstream systems don’t consume vibes. Every handoff between the model and your software needs a contract.

Reliability engineering — fallbacks, retries, timeouts, error handling. Models go down, rate limits hit, latency spikes. Production AI needs the same resilience discipline as any distributed system, because it is one.

Below the waterline: operations

This is the layer nobody demos and everybody pays for.

Cost and latency optimization — token budgets, smaller models, smart routing, caching, batching, streaming. The pilot costs $200. Production costs $200,000. Nobody budgets for it until the first invoice arrives.

CI/CD for AI — testing prompts, workflows, models, deployments. Prompts are code now: versioned, reviewed, tested, rolled back.

Production monitoring — drift, quality degradation, usage, incidents. Models don’t stay good by themselves. Data drifts, user behavior shifts, and yesterday’s eval scores quietly decay.

Security, governance, and human-in-the-loop — authentication, secrets, prompt-injection defenses, data governance, approvals and escalation workflows. Mapped to something like the NIST AI Risk Management Framework, this is the layer that keeps you out of the news.

What this means for leaders

Three implications if you’re funding or hiring for AI:

Budget for the iceberg, not the tip. If your AI budget covers models and prompts, you’ve budgeted for 10% of the system. The teams that blow up do so in the submerged 90% — usually on data and evals they never planned for.

Hire for what’s below the waterline. Anyone can demo. In interviews, skip the chatbot questions and ask about the submerged part: how they’d debug wrong answers in production, cut inference cost 10x, design an eval harness, or make AI writes to a CRM safe. The answers reveal who’s shipped and who’s watched tutorials.

Don’t let the pilot set the architecture. Two-week pilots prove the tip. That’s their job. But plan the production architecture — retrieval, evals, guardrails, cost controls — from day one, or you’ll rebuild everything the pilot validated.

The demo is the tip

None of this is an argument against demos. Demos are how you learn, get funding, and build momentum. Just don’t mistake the tip for the iceberg.

The organizations shipping AI that actually works all learned the same lesson, usually the expensive way: production is the iceberg. Build for what’s below the waterline, and the tip takes care of itself.


— Jugal

Comments

Thanks for the comment, will get back to you soon… Jugal Shah