Engineering Vibe Coding: A Five-Layer System for Production-Grade AI-Generated Code

AIAgentWorkflowClaude Code

This post is adapted from internal team practices.

When I first started using Claude Code half a year ago, my workflow looked like this: I'd dump a requirement in, the AI would churn out some code, I'd copy-paste it to run it, ask follow-ups if it broke, commit if it worked. A simple CRUD endpoint ended up as a 20-turn conversation — the code ran, but the quality was terrible: no error handling, no logging, no tests, inconsistent naming, and SQL injection vulnerabilities.

Every day I was amazed by AI, and every day I committed code worse than what I'd write myself.

After a few months, I completely engineered the team's Claude Code workflow into a five-layer system: Rules, Process, Memory, Skills, and Collaboration. Now the average team requirement produces code that goes straight into a PR from AI's first attempt. Human review time dropped from 40 minutes to 8. Here's the complete breakdown.

1. Diagnosis: Why "Raw AI" Can't Produce Production-Grade Code

Before diving into solutions, let's look at the typical "raw AI" failure modes.

After two weeks observing the team across 50 requirement interactions, problems clustered into five categories:

Problem TypeShareTypical Manifestation
Context loss38%Changed module A but forgot module B's corresponding reference; interface fields misaligned
Style drift22%Naming flips between camelCase and snake_case; log formats inconsistent
Process skipping18%Skips tests and delivers directly; commits without verifying
Repeated mistakes14%Same anti-patterns appear repeatedly (empty catch, hardcoded URLs, NPE ignored)
Scope creep8%Asked to modify one method, rewrites half the file unprompted

All five have a common root: they are engineering problems, not AI capability problems. AI can write good code — it just needs clear boundaries, context, and process constraints. Raw usage has all of these missing. You're treating AI as a "search engine that writes code," not an engineer that needs to be managed.

Once that's clear, all the engineering work follows naturally.

2. Layer One: Rules — Using CLAUDE.md to Set Boundaries for AI

CLAUDE.md is Claude Code's project-level config file, placed at the repo root and loaded automatically at the start of every session. This is the first gate of AI engineering.

Most teams' CLAUDE.md looks like this:

markdown
# Project Info
This is a Spring Boot project using Java 17.

This is effectively blank — AI learns what tech stack is used, but not how to use it. A working CLAUDE.md needs three things: boundaries, rules, and conventions.

The key principle: write what AI trips up on, not what AI already knows. Don't write "Spring Boot is a Java framework" — write "@Transactional must include rollbackFor" — the nuance that's easy to miss.

Another common mistake: CLAUDE.md that's too long. I've seen 800-line CLAUDE.md files. The result is AI losing the plot. Keep it under 300 lines, use short clause-style sentences, not paragraphs. If content is genuinely heavy, split it across files and reference them from CLAUDE.md.

3. Layer Two: Process — The Nine-Step Method to Prevent AI Cutting Corners

Rules alone aren't enough. When AI receives a requirement, its default behavior is "produce code as fast as possible," skipping whatever it can. You need a process that forces AI through: requirement confirmation → solution design → coding → self-review → testing → delivery.

The team's nine-step method emerged from lived experience:

plaintext
1. Requirement confirmation → AI must restate the requirement and list edge cases / questions
2. PRD drafting → Requirements of medium+ complexity must produce a PRD, saved to disk
3. Solution design → Product / UI / tech, each reviewed independently
4. Incremental implementation → Implement by subtask, report after each one
5. Self-review → Must complete review checklist before entering testing
6. Test verification → Unit tests + integration tests + manual test matrix
7. Delivery acceptance → Output delivery checklist (changed files, SQL, rollback plan)
8. Deployment guide → Output deployment steps, rollback plan
9. Record archival → Auto-generate dev-log

The nine-step method's biggest value isn't getting AI to do more — it's forcing AI to expose intermediate state. With raw AI, you only see the final code; you have no visibility into what it was thinking or what it skipped. The nine-step method forces documented output at every stage, and those documents are the handles for review.

4. Layer Three: Memory — Making AI Consistent Across Sessions

AI defaults to "goldfish memory" — every conversation starts fresh. This causes two problems: best practices AI told you last time are forgotten next time; anti-patterns you corrected once reappear next time.

The solution is to crystallize "project knowledge" so AI reads it every time. This has two parts:

Project Knowledge Base — documents under docs/ that AI must read during requirement confirmation. Use Glob for precise matching, then Grep for targeted search. Ban direct global code search.

Memory Files — AI-exclusive memory under .claude/memory/, capturing cross-session feedback. Every time AI receives a correction,沉淀 it proactively. After half a year, the team has accumulated dozens of these "pitfall memories," cutting AI's repeated mistake rate from 14% to 3%.

Memory files shouldn't grow unbounded. Keep individual files under 100 lines, with concrete "why" and "how to apply" sections, not abstract principles.

5. Layer Four: Skills — Packaging Reusable Capabilities as Prompts

After a month, you'll notice AI repeatedly needs the same "hints" in certain scenarios. Every time it writes SQL, you have to remind it "use prepared statements, add LIMIT, add indexes." Every time it writes concurrent code, you remind it "use ThreadPoolExecutor, not Executors, add timeouts."

These recurring reminders are the perfect candidates for Skills.

A Skill is a pre-set prompt template invoked via the Skill tool. Compared to cramming all rules into CLAUDE.md, Skills are on-demand — lower token cost and version-controllable. A Skill file is just a markdown in the git repo, so git blame tells you exactly who added what rule, when, and why.

Skills aren't "more is better." I once saw a team with 50+ Skills. AI had to consider dozens of Skills for applicability on every task, responses slowed, and conflicts between Skills occasionally emerged. Skills should be carefully curated — one per domain is enough. When overlap is found, merge.

6. Layer Five: Collaboration — Subagents for Independent Tasks

At this layer, you're no longer "using one AI to write code" — you're orchestrating multiple AIs.

Claude Code Subagents let the main agent break complex tasks into subtasks and dispatch them to different subagents in parallel — like a team lead assigning work to engineers rather than doing everything solo.

The canonical use case: large-scale refactors. For example, migrating the entire project's logging framework from Log4j 1.x to Log4j 2 across 200+ files. The right approach uses subagents:

plaintext
Main agent (architect role):
  1. Scan the project and group files by module (50 files/group)
  2. Dispatch 5 subagents (executor role), each responsible for one group
  3. Subagents return results on completion
  4. Main agent aggregates and does cross-consistency checks
  5. Main agent handles entry-point config changes itself (only 2-3 files)

Three subagent pitfalls:

  • Pitfall 1: Subagents don't share information. If subagent A changes a shared class, subagent B doesn't know, leading to inconsistency. Solution: subagents cannot modify shared classes — those changes go through the main agent.
  • Pitfall 2: Task granularity is hard to get right. Too fine-grained,调度 overhead dominates; too coarse, isolation breaks down. The经验值: no more than 30 minutes or 50 files per subagent.
  • Pitfall 3: Subagent results need verification. "I finished" ≠ "I finished correctly." The main agent must spot-check — randomly verify 20% of subagent outputs manually or via tooling.

7. Real Results from All Five Layers Together

After completing all five layers, here's what the team's numbers looked like (3 months, 80 requirements):

MetricBefore EngineeringAfter EngineeringChange
Avg. delivery time4.2 hours1.8 hours-57%
Avg. AI conversation turns186-67%
First-review pass rate22%71%+213%
Avg. rework rounds2.30.4-83%
Production bugs triggered71-86%

The numbers matter less than the logic behind them: each layer fills a specific gap in AI's default behavior.

  • Rules layer fills "AI doesn't know project conventions"
  • Process layer fills "AI cuts corners"
  • Memory layer fills "AI is inconsistent across sessions"
  • Skills layer fills "AI repeatedly makes the same mistakes"
  • Collaboration layer fills "AI's context explodes on large tasks"

Every layer is necessary. Missing one leaves a systemic漏洞.

8. Anti-Patterns to Avoid

Anti-pattern 1: Using CLAUDE.md as a documentation hub

Cramming all project docs into CLAUDE.md — 3000 lines — and AI loses the plot. CLAUDE.md should be an index pointing to other documents, not the documents themselves.

Anti-pattern 2: One-size-fits-all nine-step method

Running the full nine steps for a "change this copy / fix this typo" request is pure waste. The nine-step method must scale with complexity: simple requests simplify to three steps (confirm → code → self-review), medium ones go full process, complex ones detail every step.

Anti-pattern 3: Skill abuse

Making a Skill for every nuance, causing Skill count to explode. Skills should be for high-frequency + easy-to-mess-up rule sets. Low-frequency stuff lives in CLAUDE.md.

Anti-pattern 4: Subagent universalism

Dispatching every task to subagents turns the main agent into a dispatcher, stripping it of overall code perspective. Subagents fit "high repetition, independent context" tasks; creative tasks need the main agent doing the work.

Anti-pattern 5: Never updating memory

Writing a memory file once and forgetting about it. Anti-patterns from six months ago may be outdated now. Do a memory cleanup every quarter — delete stale entries, merge duplicates, add new ones.

Conclusion: AI Doesn't Replace Engineers — It Amplifies Them

After building this system, here's the counter-intuitive finding: teams that use AI well are actually more exhausted — but at higher-level things.

Before, exhaustion came from writing CRUD, fixing bugs, filling in tests — all handled by AI now. Now exhaustion comes from writing rules, designing processes, reviewing code, making architecture decisions — things AI can't do. The role shifted from "code producer" to "AI supervisor + architect." The work level-upped.

That's what Vibe Coding engineering really is: not replacing engineers with AI, but giving engineers time for what matters more. Rules, process, memory, skills, collaboration — these five layers aren't built for AI. They're built for your future self and the whole team.

If your team is still "raw-using AI," start today with a solid CLAUDE.md. It's the easiest step, and the highest ROI. The other four layers, build gradually. Results in three months, shaped in six, transformed in a year.

References

Comments

Sign in to leave a comment Sign in

Loading…

Back to blog