System Overview

Code Mint is an automated code generation system that transforms Jira tickets into Bitbucket pull requests using Claude as the code generation engine.

FastAPI
Async webhook server with background task processing
Claude CLI
Code generation via subprocess with session continuity
Git + BB
Worktree isolation, branch management, PR creation
graph TB
    subgraph External["External Services"]
        Jira["Jira Cloud"]
        BB["Bitbucket Cloud"]
        Conf["Confluence"]
    end
    subgraph CM["Code Mint (K8s Pod)"]
        WS["Webhook Server\n(FastAPI)"]
        O["Orchestrator"]
        C["Claude CLI\n(subprocess)"]
        G["Git Manager"]
        D["Dashboard"]
    end
    subgraph S["Persistent Storage (EFS)"]
        Jobs["Job State"]
        Learn["Learning Store"]
        Repos["Repo Clones"]
        Sess["Sessions"]
    end
    Jira -->|webhook| WS
    BB -->|PR webhook| WS
    WS --> O
    WS --> D
    O --> C
    O --> G
    O -->|comments| Jira
    G -->|push, PR| BB
    O -->|fetch| Conf
    O --> Jobs
    O --> Learn
    G --> Repos
    C --> Sess
                    

Component Architecture

21 production modules organized in 5 layers.

CORE Core Pipeline

orchestrator.py
Main workflow engine. Coordinates the full ticket-to-PR pipeline across all phases.
~3,378 LOC
prompts.py
Centralized prompt builder. All Claude CLI phase prompts. Pure functions, no side effects.
~941 LOC
context_gatherer.py
Ticket enrichment. Fetches epic, linked issues, Confluence, Figma, comments.
repo_router.py
Repository targeting. Labels → components → keywords → LLM fallback.

AGENT Agent Layer

agent.py
Plan extraction, session reading, Claude API fallback
code_validator.py
Build/test validation. Auto-detects project type.
smart_repo_matcher.py
LLM-powered repository relevance scoring

I/O External I/O

jira_client.py
Jira REST API: comments, labels, issues, links
bitbucket_client.py
Bitbucket API: PR creation and querying
git_utils.py
Git operations: clone, worktree, branch, commit, push

STATE State Management

job_store.py
Job state, phases, tokens, PR events
learning_store.py
Implementation records, agent-searchable artifacts
session_store.py
Claude session IDs for --resume continuity
plan_file_store.py
Plan markdown file management
db.py
PostgreSQL ORM (SQLAlchemy 2.0 async)
job_logger.py
Per-job structured logging + token tracking

SHARED Shared Infrastructure

models.py
Pydantic data models
config.py
Configuration loading
ticket_parser.py
Webhook normalization
session_parser.py
JSONL transcript parsing
Dependency Rules
  • orchestrator is the only module that touches all components
  • prompts is a pure function module — no side effects, no state
  • External I/O modules have no cross-dependencies
  • State stores are independent singletons via get_*() factories

Processing Pipeline

Deterministic phase pipeline with early exit points at each stage.

Webhook
~50ms
→
Pre-Processing
dedup, eligibility, rate limit
→
Context Gathering
2-5s
Repo Analysis
~50ms
→
Clone Repos
2-30s (worktree)
→
Clarification
1-2min (LLM)
Plan Mode
3-8min (optional)
→
Code Generation
2-8min (Claude CLI)
→
Review
1-5min (optional)
Commit & Push
3-5s
→
Create PR
1-2s
→
Done
Post Jira + Learn
Total Processing Time
4-10 min
Without plan mode
7-18 min
With plan + review

Session Architecture

One Claude conversation per ticket. Phases share context via --resume session_id.

sequenceDiagram
    participant O as Orchestrator
    participant C as Claude CLI
    participant S as Session Store

    Note over O,S: Phase: Clarification
    O->>C: claude --permission-mode plan
    C-->>O: Analysis + questions (session_id=abc)
    O->>S: save_session(ticket, abc)

    Note over O,S: Phase: Plan Generation
    O->>S: get_session(ticket) → abc
    O->>C: claude --resume abc --permission-mode plan
    C-->>O: Implementation plan

    Note over O,S: Phase: Code Generation
    O->>S: get_session(ticket) → abc
    O->>C: claude --resume abc --dangerously-skip-permissions
    C-->>O: Generated code + JSON output

    Note over O,S: Phase: Independent Review (NEW session)
    O->>C: claude --permission-mode plan (fresh eyes)
    C-->>O: ReviewVerdict (no prior context)
                    
Why this matters: Code generation has full memory of clarification Q&A and planning decisions. No context is repeated. The independent review is intentionally a new session — fresh eyes catch what the author doesn't.

Deployment Architecture

Horizontally scalable deployment on AWS EKS with EFS shared storage and PostgreSQL coordination.

graph TB
    subgraph Internet
        JC["Jira Cloud"]
        BC["Bitbucket Cloud"]
    end
    subgraph AWS
        subgraph EKS["EKS Cluster"]
            Ing["Ingress / Kong"]
            subgraph Pod["Code Mint Pod (N replicas)"]
                App["FastAPI\nPython 3.11"]
                CLI["Claude CLI\nNode.js 20"]
                RT["Build Runtimes\nJava 17 / Go 1.22"]
            end
        end
        subgraph EFS
            WD["workdir (50Gi)\nrepos, jobs, learning"]
            CH["claude-home (10Gi)\nCLI config, sessions"]
        end
        PG["PostgreSQL\n(optional)"]
    end
    Claude["Anthropic\nClaude API"]

    JC -->|webhook| Ing
    BC -->|webhook| Ing
    Ing --> App
    App --> CLI
    CLI --> Claude
    App --> WD
    CLI --> CH
    App -.-> PG
    App -->|push, PR| BC
    App -->|comments| JC
                    

Container Runtimes

Python 3.11 Application server
Node.js 20 Claude CLI + JS validation
Java 17 + Maven Java build/test
Go 1.22 Go build/test

Key Design Choices

  • Horizontally scalable — PostgreSQL dedup locks
  • Non-root — uid 1000 (automation user)
  • HOME override — EFS for session persistence
  • No ANTHROPIC_API_KEY — Max plan billing

Data Architecture

PostgreSQL for job and PR event tracking. File-based storage for sessions, plans, and learning.

graph LR
    subgraph PG["PostgreSQL (DATABASE_URL required)"]
        PJ["jobs"]
        PL["processing_locks"]
        PLR["learning_metadata"]
        PP["pr_events"]
    end
    subgraph Files["File-Based Stores"]
        S["sessions/*.json"]
        P["plans/*.md"]
        L["learning/records/\nlearning/index.jsonl"]
        LG["job_logs/*.log"]
    end
                    

File System Layout

{WORK_DIR}/
├── sessions/                # Claude session continuity
│   └── {TICKET-ID}.json
├── plans/                   # Implementation plans
│   └── {TICKET-ID}_v{N}.md
├── job_logs/                # Per-job log files
│   └── {JOB-ID}.log
├── learning/                # Continuous learning
│   ├── index.jsonl          # Grep-searchable index
│   └── records/
│       ├── {TICKET}.json    # Structured record
│       └── {TICKET}/
│           └── outcome.md   # Agent-readable markdown
├── {repo}/                  # Main clone (shared)
└── {repo}-worktrees/        # Per-ticket isolation
    └── {TICKET}/

Integration Architecture

External service dependencies and authentication methods.

Jira REST API

Basic Auth (email + API token)
  • GET /issue/{id} Epic, links
  • GET /issue/{id}/comment Comments
  • GET /issue/{id}/remotelink Confluence/Figma
  • POST /issue/{id}/comment Post results
  • PUT /issue/{id} Add labels

Bitbucket REST API

Bearer token (preferred) or Basic Auth
  • POST /pullrequests Create PR
  • GET /pullrequests?q= Find existing
  • git clone/push HTTPS Basic Auth

Claude CLI

OAuth token (CLAUDE_CODE_OAUTH_TOKEN)
  • --permission-mode plan Read-only phases
  • --dangerously-skip-permissions Code gen
  • --resume {id} Session continuity
  • --add-dir Learning, context, repos

Confluence

Basic Auth (optional)
  • GET /content/{id} Fetch page body
  • Used during context gathering phase

Concurrency & Isolation

Git worktrees provide per-ticket isolation. Deduplication prevents duplicate processing.

Worktree Strategy

graph TB
    MC["Main Clone\n{repo}/.git\n(shared, never modified)"]
    W1["{repo}-worktrees/PROJ-123/"]
    W2["{repo}-worktrees/PROJ-456/"]
    W3["{repo}-worktrees/PROJ-789/"]
    MC -->|"git worktree add"| W1
    MC -->|"git worktree add"| W2
    MC -->|"git worktree add"| W3
                    

Deduplication Flow

flowchart TD
    WH["Webhook"] --> ACL{"Auto-code\nlabel check"}
    ACL -->|missing| REJ["Reject\n(skip dedup cache)"]
    ACL -->|ok| PG{"PostgreSQL?"}
    PG -->|yes| LOCK["try_acquire_lock\n(10s window)"]
    PG -->|no| MEM["In-memory cache\n(10s window)"]
    LOCK -->|acquired| RUN{"Running job?"}
    LOCK -->|not acquired| DUP["SKIP: duplicate"]
    MEM -->|unique| RUN
    MEM -->|duplicate| DUP
    RUN -->|yes| SKIP["SKIP: already running"]
    RUN -->|no| GO["Continue"]
                    

Learning Feedback Loop

Each successful PR improves future code generation through recorded learnings.

graph TB
    subgraph Gen["Code Generation"]
        P["Prompt Builder"]
        C["Claude CLI"]
        S["Agent searches\nlearning/ via grep"]
    end
    subgraph Out["Output"]
        PR["Bitbucket PR"]
        JSON["CodeGenOutput\nsummary, files, decisions"]
    end
    subgraph Learn["Learning Store"]
        Rec["JSON Record"]
        MD["outcome.md"]
        Idx["index.jsonl"]
    end

    P --> C
    S --> C
    C --> PR
    C --> JSON
    JSON --> Rec
    Rec --> MD
    Rec --> Idx
    Idx -->|"grep similar"| S
    MD -->|"cat details"| S
                    
How agents use learnings:
  1. Learning directory injected via --add-dir learning/
  2. Agent runs grep "payment" learning/index.jsonl
  3. Agent reads cat learning/records/PROJ-123/outcome.md
  4. Past patterns inform current code generation

Architecture Decisions

Key decisions with rationale, emerged from 40 phases of evolution.

Single Session over Parallel Subagents

~1,500 LOC deleted

Replaced multiple parallel Claude subagents with single-session processing. Simpler and equally effective.

Lesson: Complexity is not sophistication.

Agent-Searchable Files over Pre-Injected Context

3 modules deleted

Mount directories via --add-dir instead of classifying and pre-injecting context. Agents find more relevant context by searching interactively.

Lesson: Give agents tools instead of pre-chewing their food.

JSON Structured Output over Marker Parsing

Switched from regex-parsed ### FILE: markers to CodeGenOutput JSON model validated by Pydantic.

Lesson: Parse only what the system needs for routing.

Independent Review over Self-Review

Separate read-only Claude session with no memory of generation. Fresh eyes catch what the author doesn't.

Lesson: Verification must be structurally independent from creation.

LLM-Powered Clarification over Static Rules

LLM reasons about the ticket in context of the target repo. Better questions, fewer false alarms.

Lesson: If the decision requires judgment, use a model that can exercise judgment.

Centralized Prompts over Scattered Strings

All prompts in prompts.py. Single source of truth for the system's personality.

Lesson: Prompts are code. Treat them with the same discipline.

Multi-Tag Classification over Single Enum

LLM generates tags: list[str] instead of single classification. A ticket can be both payment and API.

Lesson: Don't force false choices in classification.

PostgreSQL as Required Store

PostgreSQL for jobs, PR events, and processing locks. Row-level locking for multi-pod dedup and concurrency control.

Lesson: Remove dual-path complexity — one store, one code path.

Git Worktrees for Isolation

Each ticket gets an isolated worktree from a shared main clone. No branch switching, no conflicts.

Lesson: Filesystem isolation beats branch-based isolation.

Design Principles

Emerged from 40 phases of evolution.

1
Remove logic, add tools
Strip orchestrator-driven decision-making. Give agents tools and let them decide.
2
Agent-searchable over pre-injected
Mount files, don't stuff prompts. Let agents explore what they need.
3
Single code path
One flow for single-repo and multi-repo. No mode switches.
4
Fail open, log everything
Non-critical failures log warnings but don't block the pipeline.
5
Structural independence
Verification must be separate from creation. Fresh eyes catch what authors miss.
6
Session continuity
One conversation per ticket from clarification through code generation.
7
File-based simplicity
Raw markdown for plans, JSONL for indexes, JSON for state. No custom formats.