System Overview
Code Mint is an automated code generation system that transforms Jira tickets into Bitbucket pull requests using Claude as the code generation engine.
graph TB
subgraph External["External Services"]
Jira["Jira Cloud"]
BB["Bitbucket Cloud"]
Conf["Confluence"]
end
subgraph CM["Code Mint (K8s Pod)"]
WS["Webhook Server\n(FastAPI)"]
O["Orchestrator"]
C["Claude CLI\n(subprocess)"]
G["Git Manager"]
D["Dashboard"]
end
subgraph S["Persistent Storage (EFS)"]
Jobs["Job State"]
Learn["Learning Store"]
Repos["Repo Clones"]
Sess["Sessions"]
end
Jira -->|webhook| WS
BB -->|PR webhook| WS
WS --> O
WS --> D
O --> C
O --> G
O -->|comments| Jira
G -->|push, PR| BB
O -->|fetch| Conf
O --> Jobs
O --> Learn
G --> Repos
C --> Sess
Component Architecture
21 production modules organized in 5 layers.
CORE Core Pipeline
AGENT Agent Layer
I/O External I/O
STATE State Management
SHARED Shared Infrastructure
- orchestrator is the only module that touches all components
- prompts is a pure function module — no side effects, no state
- External I/O modules have no cross-dependencies
- State stores are independent singletons via
get_*()factories
Processing Pipeline
Deterministic phase pipeline with early exit points at each stage.
Session Architecture
One Claude conversation per ticket. Phases share context via --resume session_id.
sequenceDiagram
participant O as Orchestrator
participant C as Claude CLI
participant S as Session Store
Note over O,S: Phase: Clarification
O->>C: claude --permission-mode plan
C-->>O: Analysis + questions (session_id=abc)
O->>S: save_session(ticket, abc)
Note over O,S: Phase: Plan Generation
O->>S: get_session(ticket) → abc
O->>C: claude --resume abc --permission-mode plan
C-->>O: Implementation plan
Note over O,S: Phase: Code Generation
O->>S: get_session(ticket) → abc
O->>C: claude --resume abc --dangerously-skip-permissions
C-->>O: Generated code + JSON output
Note over O,S: Phase: Independent Review (NEW session)
O->>C: claude --permission-mode plan (fresh eyes)
C-->>O: ReviewVerdict (no prior context)
Deployment Architecture
Horizontally scalable deployment on AWS EKS with EFS shared storage and PostgreSQL coordination.
graph TB
subgraph Internet
JC["Jira Cloud"]
BC["Bitbucket Cloud"]
end
subgraph AWS
subgraph EKS["EKS Cluster"]
Ing["Ingress / Kong"]
subgraph Pod["Code Mint Pod (N replicas)"]
App["FastAPI\nPython 3.11"]
CLI["Claude CLI\nNode.js 20"]
RT["Build Runtimes\nJava 17 / Go 1.22"]
end
end
subgraph EFS
WD["workdir (50Gi)\nrepos, jobs, learning"]
CH["claude-home (10Gi)\nCLI config, sessions"]
end
PG["PostgreSQL\n(optional)"]
end
Claude["Anthropic\nClaude API"]
JC -->|webhook| Ing
BC -->|webhook| Ing
Ing --> App
App --> CLI
CLI --> Claude
App --> WD
CLI --> CH
App -.-> PG
App -->|push, PR| BC
App -->|comments| JC
Container Runtimes
| Python 3.11 | Application server |
| Node.js 20 | Claude CLI + JS validation |
| Java 17 + Maven | Java build/test |
| Go 1.22 | Go build/test |
Key Design Choices
- Horizontally scalable — PostgreSQL dedup locks
- Non-root — uid 1000 (automation user)
- HOME override — EFS for session persistence
- No ANTHROPIC_API_KEY — Max plan billing
Data Architecture
PostgreSQL for job and PR event tracking. File-based storage for sessions, plans, and learning.
graph LR
subgraph PG["PostgreSQL (DATABASE_URL required)"]
PJ["jobs"]
PL["processing_locks"]
PLR["learning_metadata"]
PP["pr_events"]
end
subgraph Files["File-Based Stores"]
S["sessions/*.json"]
P["plans/*.md"]
L["learning/records/\nlearning/index.jsonl"]
LG["job_logs/*.log"]
end
File System Layout
{WORK_DIR}/
├── sessions/ # Claude session continuity
│ └── {TICKET-ID}.json
├── plans/ # Implementation plans
│ └── {TICKET-ID}_v{N}.md
├── job_logs/ # Per-job log files
│ └── {JOB-ID}.log
├── learning/ # Continuous learning
│ ├── index.jsonl # Grep-searchable index
│ └── records/
│ ├── {TICKET}.json # Structured record
│ └── {TICKET}/
│ └── outcome.md # Agent-readable markdown
├── {repo}/ # Main clone (shared)
└── {repo}-worktrees/ # Per-ticket isolation
└── {TICKET}/
Integration Architecture
External service dependencies and authentication methods.
Jira REST API
GET /issue/{id}Epic, linksGET /issue/{id}/commentCommentsGET /issue/{id}/remotelinkConfluence/FigmaPOST /issue/{id}/commentPost resultsPUT /issue/{id}Add labels
Bitbucket REST API
POST /pullrequestsCreate PRGET /pullrequests?q=Find existinggit clone/pushHTTPS Basic Auth
Claude CLI
--permission-mode planRead-only phases--dangerously-skip-permissionsCode gen--resume {id}Session continuity--add-dirLearning, context, repos
Confluence
GET /content/{id}Fetch page body- Used during context gathering phase
Concurrency & Isolation
Git worktrees provide per-ticket isolation. Deduplication prevents duplicate processing.
Worktree Strategy
graph TB
MC["Main Clone\n{repo}/.git\n(shared, never modified)"]
W1["{repo}-worktrees/PROJ-123/"]
W2["{repo}-worktrees/PROJ-456/"]
W3["{repo}-worktrees/PROJ-789/"]
MC -->|"git worktree add"| W1
MC -->|"git worktree add"| W2
MC -->|"git worktree add"| W3
Deduplication Flow
flowchart TD
WH["Webhook"] --> ACL{"Auto-code\nlabel check"}
ACL -->|missing| REJ["Reject\n(skip dedup cache)"]
ACL -->|ok| PG{"PostgreSQL?"}
PG -->|yes| LOCK["try_acquire_lock\n(10s window)"]
PG -->|no| MEM["In-memory cache\n(10s window)"]
LOCK -->|acquired| RUN{"Running job?"}
LOCK -->|not acquired| DUP["SKIP: duplicate"]
MEM -->|unique| RUN
MEM -->|duplicate| DUP
RUN -->|yes| SKIP["SKIP: already running"]
RUN -->|no| GO["Continue"]
Learning Feedback Loop
Each successful PR improves future code generation through recorded learnings.
graph TB
subgraph Gen["Code Generation"]
P["Prompt Builder"]
C["Claude CLI"]
S["Agent searches\nlearning/ via grep"]
end
subgraph Out["Output"]
PR["Bitbucket PR"]
JSON["CodeGenOutput\nsummary, files, decisions"]
end
subgraph Learn["Learning Store"]
Rec["JSON Record"]
MD["outcome.md"]
Idx["index.jsonl"]
end
P --> C
S --> C
C --> PR
C --> JSON
JSON --> Rec
Rec --> MD
Rec --> Idx
Idx -->|"grep similar"| S
MD -->|"cat details"| S
- Learning directory injected via
--add-dir learning/ - Agent runs
grep "payment" learning/index.jsonl - Agent reads
cat learning/records/PROJ-123/outcome.md - Past patterns inform current code generation
Architecture Decisions
Key decisions with rationale, emerged from 40 phases of evolution.
Single Session over Parallel Subagents
Replaced multiple parallel Claude subagents with single-session processing. Simpler and equally effective.
Agent-Searchable Files over Pre-Injected Context
Mount directories via --add-dir instead of classifying and pre-injecting context. Agents find more relevant context by searching interactively.
JSON Structured Output over Marker Parsing
Switched from regex-parsed ### FILE: markers to CodeGenOutput JSON model validated by Pydantic.
Independent Review over Self-Review
Separate read-only Claude session with no memory of generation. Fresh eyes catch what the author doesn't.
LLM-Powered Clarification over Static Rules
LLM reasons about the ticket in context of the target repo. Better questions, fewer false alarms.
Centralized Prompts over Scattered Strings
All prompts in prompts.py. Single source of truth for the system's personality.
Multi-Tag Classification over Single Enum
LLM generates tags: list[str] instead of single classification. A ticket can be both payment and API.
PostgreSQL as Required Store
PostgreSQL for jobs, PR events, and processing locks. Row-level locking for multi-pod dedup and concurrency control.
Git Worktrees for Isolation
Each ticket gets an isolated worktree from a shared main clone. No branch switching, no conflicts.
Design Principles
Emerged from 40 phases of evolution.