first pass at the newspaper builder
Test / test (push) Has been cancelled

This commit is contained in:
2026-09-14 11:57:22 +10:00
commit bec1eaac87
497 changed files with 178953 additions and 0 deletions
@@ -0,0 +1,146 @@
# Architecture Decision Records (ADRs)
## Purpose
ADRs capture significant architectural decisions, the context that drove them, and the consequences of choosing one option over alternatives. They create a decision journal that explains WHY the architecture looks the way it does.
## When to Create an ADR
Create an ADR when the decision:
- Is difficult to reverse once implemented (database choice, API contract, framework)
- Affects multiple teams or components
- Has significant cost, performance, or security implications
- Was debated — if there was disagreement, the reasoning needs documentation
- Changes a previous architectural direction
Do NOT create an ADR for:
- Routine implementation choices (variable names, code formatting)
- Decisions that are trivially reversible
- Standard practices already documented in team guidelines
## ADR Template
```markdown
# ADR-[NUMBER]: [Title - Short Descriptive Name]
## Status
[Proposed | Accepted | Deprecated | Superseded by ADR-XXX]
## Date
[YYYY-MM-DD]
## Context
What is the issue or situation that motivates this decision? Describe the forces
at play: technical constraints, business requirements, team capabilities, timeline
pressures, and any other factors influencing the decision.
Be specific. Include:
- What problem are we solving?
- What constraints exist (budget, timeline, team skill, compliance)?
- What quality attributes matter most (performance, security, maintainability)?
- What existing decisions or systems does this interact with?
## Decision
State the decision clearly and concisely. Use active voice:
"We will use PostgreSQL as the primary data store for the order management service."
Not: "It was decided that PostgreSQL might be a good option."
## Consequences
### Positive
- What becomes easier, faster, or better as a result of this decision?
- What risks are mitigated?
### Negative
- What becomes harder or more expensive?
- What new risks are introduced?
- What capabilities are foreclosed?
### Neutral
- What trade-offs are we accepting?
- What follow-up decisions will be needed?
## Alternatives Considered
### Alternative 1: [Name]
- Description: Brief explanation of this option
- Pros: What would this have given us?
- Cons: Why did we reject it?
### Alternative 2: [Name]
- Description: Brief explanation of this option
- Pros: What would this have given us?
- Cons: Why did we reject it?
## References
- Links to relevant RFCs, design documents, benchmark results, or discussions
```
## Numbering Convention
Use sequential numbering with zero-padding:
```
ADR-001-use-postgresql-for-orders.md
ADR-002-adopt-event-driven-integration.md
ADR-003-select-react-for-frontend.md
```
### Naming Rules
- Sequential numbers, never reused (even if deprecated)
- Kebab-case descriptive suffix after the number
- Store in a dedicated `/docs/adr/` or `/architecture/decisions/` directory
- Include an `index.md` that lists all ADRs with status and one-line summary
## ADR Lifecycle
### Proposed
The decision is under discussion. The ADR is a draft open for review and feedback. Include it in pull requests or architecture review meetings.
### Accepted
The decision has been approved and should be followed. Implementation can proceed. Record the acceptance date and approver(s).
### Deprecated
The decision is no longer relevant — the system or feature it applied to has been removed. Keep the ADR for historical context but mark it clearly.
### Superseded
A new decision has replaced this one. Link to the new ADR:
```
## Status
Superseded by [ADR-015](./ADR-015-migrate-to-dynamodb.md)
```
The superseding ADR should reference the original:
```
## Context
This decision supersedes [ADR-003](./ADR-003-select-postgresql.md) because...
```
## Best Practices
### Writing Quality
- Write for a future reader who has no context — assume they joined the team after the decision was made
- Focus on WHY, not just WHAT — the code shows what was built; the ADR explains why
- Be honest about trade-offs — every decision has downsides; documenting them builds trust
- Include quantitative data when available (benchmark results, cost estimates, load test numbers)
### Process
- Create the ADR BEFORE implementation, not after — it is a decision tool, not documentation
- Review ADRs in pull requests alongside the code they influence
- Revisit ADRs quarterly — some may need updating as circumstances change
- Keep ADRs concise — one to two pages is ideal; longer suggests the scope is too broad
### Common Mistakes
- Writing ADRs after the fact as documentation (they should drive the decision)
- Omitting alternatives (makes the decision look unconsidered)
- Not recording negative consequences (creates false confidence)
- Making ADRs too granular (implementation details do not need ADRs)
- Letting ADRs become stale without status updates
## Example ADR Index
| ADR | Title | Status | Date |
|-----|-------|--------|------|
| 001 | Use PostgreSQL for order management | Accepted | 2024-01-15 |
| 002 | Adopt event-driven integration between services | Accepted | 2024-02-01 |
| 003 | Select React with TypeScript for frontend | Accepted | 2024-02-10 |
| 004 | Use JWT for API authentication | Superseded by 012 | 2024-03-01 |
| 005 | Deploy to AWS ECS Fargate | Accepted | 2024-03-15 |
@@ -0,0 +1,84 @@
# Architecture Guide
## Architectural Style Selection
Choose based on system characteristics:
| Style | When to Use | Avoid When |
|-------|-------------|------------|
| Modular Monolith | Single team, shared database, <10 bounded contexts | Independent scaling needed per module |
| Microservices | Multiple teams, independent deploy cycles, polyglot needs | Small team, simple domain, early-stage product |
| Event-Driven | Async workflows, audit trails, temporal decoupling needed | Strong consistency required everywhere |
| Serverless | Sporadic traffic, event-triggered compute, rapid prototyping | Long-running processes, predictable high throughput |
| Layered (N-Tier) | CRUD-dominant apps, well-understood domain | Complex domain logic, high-performance paths |
## Component Boundary Identification
A component boundary is correct when:
- The component has a single, nameable responsibility
- It owns its data (no shared mutable state across boundaries)
- Changes to its internals do not ripple to other components
- It can be tested in isolation with stub/mock dependencies
- It has a clear public API surface (no back-channel coupling)
Red flags for wrong boundaries:
- Two components that always deploy together
- Circular dependencies between components
- A component that is just a pass-through proxy
- Shared database tables written by multiple components
## Design Pattern Checklist
For each pattern decision, evaluate:
1. **Problem fit**: Does the pattern solve the specific problem, not a hypothetical one?
2. **Complexity cost**: Is the added indirection justified by the flexibility gained?
3. **Team familiarity**: Can the team maintain this pattern without the architect present?
4. **Testing impact**: Does the pattern make the system easier or harder to test?
Common patterns and their AI-DLC application:
- **Repository**: Data access abstraction -- use when persistence technology may change
- **CQRS**: Separate read/write models -- use when read and write patterns diverge significantly
- **Saga/Orchestrator**: Distributed transaction coordination -- use for cross-service workflows
- **Strategy**: Runtime algorithm selection -- use when behavior varies by configuration/tenant
- **Adapter/Port**: External integration isolation -- always use for third-party dependencies
## ADR Format
```
# ADR-NNN: [Decision Title]
## Status: [Proposed | Accepted | Deprecated | Superseded by ADR-NNN]
## Context
[What forces are at play? What constraints exist? What problem are we solving?]
## Decision
[What is the change we are making? Be specific and actionable.]
## Consequences
[What becomes easier? What becomes harder? What are the trade-offs?]
## Alternatives Considered
[What other options were evaluated and why were they rejected?]
```
## Infrastructure Pattern Alignment
When designing application topology, validate against infrastructure:
- **Stateless services**: Can scale horizontally behind a load balancer
- **Stateful services**: Need sticky sessions, distributed cache, or dedicated instances
- **Background workers**: Queue-driven, idempotent, with dead-letter handling
- **Scheduled jobs**: Cron-based, must handle overlapping executions
- **Edge functions**: Latency-sensitive, limited runtime, no persistent connections
## Reverse Engineering Synthesis Checklist
When receiving code scan results from Developer:
1. Identify the dominant architectural style (or lack thereof)
2. Map discovered components to bounded contexts
3. Trace data flow paths (request entry to persistence)
4. Flag coupling hotspots (high fan-in/fan-out modules)
5. Identify missing boundaries (God classes, shared state)
6. Assess test coverage alignment with architectural risk
7. Document observed patterns vs. intended patterns
8. Produce a component inventory with health ratings (healthy/at-risk/degraded)
@@ -0,0 +1,170 @@
# Architecture Patterns
## Purpose
Select the right architectural style for your system's requirements. No pattern is universally best — each encodes trade-offs between complexity, scalability, team autonomy, and operational cost.
## Pattern Overview
| Pattern | Best For | Team Size | Complexity |
|---------|----------|-----------|------------|
| Modular Monolith | Most new projects, small-medium teams | 1-4 teams | Low-Medium |
| Microservices | Large orgs, independent deployment needs | 5+ teams | High |
| Serverless | Event-driven workloads, variable traffic | Any | Medium |
| Event-Driven | Async workflows, loose coupling, audit trails | 2+ teams | Medium-High |
| CQRS | Read-heavy with complex queries, separate scaling needs | 2+ teams | High |
| Hexagonal | Testability, swappable infrastructure, long-lived systems | Any | Medium |
## Modular Monolith
### Description
A single deployable unit with well-defined internal module boundaries. Modules communicate through explicit interfaces, not shared database tables.
### When to Use
- Starting a new product (default choice)
- Team is small (fewer than 20 developers)
- Domain boundaries are not yet clear
- Deployment simplicity is valued
### Key Rules
- Each module owns its data (separate schemas or schema prefixes)
- Modules communicate via defined interfaces (not direct database queries across modules)
- Module dependencies are acyclic and explicitly declared
- Extract to microservices later when boundaries are proven
### Trade-offs
- (+) Simple deployment, debugging, and testing
- (+) Refactoring across modules is straightforward
- (-) Scaling is all-or-nothing (cannot scale one module independently)
- (-) Technology choices are shared across all modules
## Microservices
### Description
Independently deployable services, each owning a bounded context with its own data store.
### When to Use
- Multiple teams need to deploy independently
- Different parts of the system have very different scaling needs
- Polyglot technology requirements
- Organization has mature DevOps and observability practices
### Key Rules
- Each service owns its database — no shared data stores
- Services communicate via APIs or events, never direct database access
- Design for failure: every remote call can fail
- Deploy independently with backward-compatible API changes
### Trade-offs
- (+) Independent deployment, scaling, and technology choice
- (+) Team autonomy and clear ownership boundaries
- (-) Distributed system complexity (network failures, data consistency, debugging)
- (-) Operational overhead (monitoring, deployment pipelines per service)
- (-) Integration testing is difficult
## Serverless
### Description
Application logic runs in ephemeral, event-triggered functions managed by the cloud provider. No server provisioning or management.
### When to Use
- Event-driven workloads (file processing, webhooks, scheduled tasks)
- Highly variable traffic with long idle periods
- Rapid prototyping with minimal infrastructure
- Cost optimization for low or bursty traffic
### Key Rules
- Functions should be stateless — use external stores for state
- Keep cold start impact low (small bundles, provisioned concurrency for latency-sensitive paths)
- Design for idempotency — events may be delivered more than once
- Set concurrency limits to protect downstream systems
### Trade-offs
- (+) Zero infrastructure management, pay-per-use pricing
- (+) Automatic scaling to zero and to peak
- (-) Cold start latency (problematic for synchronous user-facing requests)
- (-) Vendor lock-in to cloud provider's function runtime and event sources
- (-) Debugging distributed function chains is difficult
- (-) Execution time limits (15 minutes on AWS Lambda)
## Event-Driven Architecture
### Description
Components communicate by producing and consuming events through a message broker. Producers do not know about consumers.
### When to Use
- Loose coupling between components is essential
- Workflows are asynchronous (order processing, notifications, data pipelines)
- Audit trails and event replay are needed
- Multiple consumers need to react to the same event
### Key Patterns
- **Event Notification**: Signal that something happened; consumer fetches details
- **Event-Carried State Transfer**: Event contains all data the consumer needs
- **Event Sourcing**: Store the sequence of events as the source of truth (not current state)
### Trade-offs
- (+) Loose coupling, independent scalability, natural audit trail
- (-) Eventual consistency (not all components see changes simultaneously)
- (-) Event ordering and deduplication complexity
- (-) Debugging event flows requires distributed tracing
## CQRS (Command Query Responsibility Segregation)
### Description
Separate the write model (commands that change state) from the read model (queries that return data). Each can use different data stores and schemas optimized for their purpose.
### When to Use
- Read and write patterns are very different (read-heavy with complex projections)
- Read and write sides need to scale independently
- Domain model is complex and queries require flattened/denormalized views
### Key Rules
- Commands validate business rules and write to the write store
- Events propagate changes to the read store (eventually consistent)
- Read models are disposable — they can be rebuilt from the event stream
- Accept eventual consistency between write and read sides
### Trade-offs
- (+) Optimized read and write performance independently
- (+) Read models tailored to specific UI needs without compromising domain model
- (-) Increased complexity (two models, synchronization, eventual consistency)
- (-) Overkill for simple CRUD applications
## Hexagonal Architecture (Ports and Adapters)
### Description
Business logic at the center, surrounded by ports (interfaces) and adapters (implementations). External systems connect through adapters; business logic never depends on infrastructure.
### Structure
```
Adapters (Infrastructure)
└── Ports (Interfaces)
└── Domain (Business Logic) ← depends on nothing external
```
### When to Use
- Business logic is complex and must be testable in isolation
- Infrastructure may change (swap database, replace message broker, change cloud provider)
- Long-lived systems where technology choices will evolve
### Key Rules
- Domain layer has zero dependencies on frameworks or infrastructure
- All external interactions go through port interfaces defined by the domain
- Adapters implement ports and translate between domain and external formats
- Tests can use in-memory adapters (no database, no network required)
## Migration Paths
| From | To | Approach |
|------|----|----------|
| Monolith | Modular Monolith | Identify module boundaries, enforce interface contracts, separate data |
| Modular Monolith | Microservices | Extract one module at a time (strangler fig pattern), starting with the least coupled |
| Monolith | Serverless | Extract event-driven workflows first (background jobs, file processing) |
| Microservices | Modular Monolith | Consolidate when services are too small, team boundaries have shifted, or operational cost exceeds benefit |
## Decision Framework
1. Start with a modular monolith unless you have proven reasons not to
2. Extract to microservices only when team autonomy or independent scaling demands it
3. Use serverless for event-driven workloads and glue code, not for core request-response APIs
4. Apply CQRS only when read and write models genuinely diverge
5. Use hexagonal architecture in the domain layer regardless of the outer architecture style
@@ -0,0 +1,134 @@
# Domain-Driven Design Patterns
## Purpose
DDD aligns software design with business reality. It provides patterns for modeling complex domains so that the code structure mirrors how the business thinks and operates.
## Strategic Design
### Bounded Contexts
A bounded context is an explicit boundary within which a domain model exists. The same real-world concept (e.g., "Customer") can have different meanings in different contexts.
**Example**:
- Sales context: Customer has leads, deals, revenue potential
- Support context: Customer has tickets, SLAs, satisfaction score
- Billing context: Customer has invoices, payment methods, credit limit
**Rules**:
- Each bounded context owns its data and logic — no shared database between contexts
- Communication between contexts uses well-defined interfaces (APIs, events)
- Each context can use different technology stacks and deployment strategies
### Context Mapping Patterns
| Pattern | Relationship | Use When |
|---------|-------------|----------|
| **Shared Kernel** | Two teams share a small common model | Teams are closely aligned and can coordinate changes |
| **Customer-Supplier** | Upstream supplies data, downstream consumes | Clear provider/consumer relationship between teams |
| **Conformist** | Downstream adopts upstream's model as-is | Upstream has no incentive to accommodate downstream needs |
| **Anti-Corruption Layer** | Downstream translates upstream's model | Integrating with legacy systems or external services |
| **Open Host Service** | Upstream publishes a well-defined protocol | Multiple consumers need access to a context's capabilities |
| **Published Language** | Shared interchange format (e.g., JSON schema) | Cross-context communication needs a stable contract |
| **Separate Ways** | No integration | The cost of integration exceeds the benefit |
### Anti-Corruption Layer (ACL)
When integrating with an external or legacy system whose model does not match yours:
```
Your Domain Model <-> ACL (translates) <-> External System Model
```
The ACL isolates your clean model from the external system's concepts. If the external system changes, only the ACL needs updating.
## Tactical Design
### Entities
Objects defined by their identity, not their attributes. Two customers with the same name are different customers if they have different IDs.
- Have a unique identifier that persists across state changes
- Mutable — their state changes over time
- Example: User, Order, Account
### Value Objects
Objects defined by their attributes, not their identity. Two Money objects with the same amount and currency are interchangeable.
- Immutable — create a new instance instead of modifying
- No identity — equality is based on all attributes
- Example: EmailAddress, Money, DateRange, Address
- Prefer value objects over primitives (an email is not just a string)
### Aggregates
A cluster of entities and value objects treated as a single unit for data changes.
**Rules**:
- Each aggregate has exactly one **aggregate root** — the only entity external code can reference
- All changes to the aggregate go through the root (enforces invariants)
- Transactions should not span multiple aggregates
- Reference other aggregates by ID, not by object reference
- Keep aggregates small — large aggregates cause contention and performance issues
**Example**:
```
Order (Aggregate Root)
├── OrderLine (Entity, only accessible through Order)
├── ShippingAddress (Value Object)
└── OrderStatus (Value Object)
```
External code calls `order.addItem(product, quantity)` — never modifies OrderLine directly.
### Domain Events
Record something meaningful that happened in the domain. Events are past-tense facts.
**Naming**: `[Entity][PastTenseVerb]` — OrderPlaced, PaymentReceived, InventoryDepleted
**Structure**:
- Event ID (unique)
- Timestamp (when it occurred)
- Aggregate ID (which aggregate produced it)
- Payload (relevant data at the time of the event)
**Uses**:
- Trigger side effects in other bounded contexts (eventual consistency)
- Build audit trails and event sourcing
- Enable loose coupling — publishers do not know about subscribers
### Repository Pattern
Provides a collection-like interface for accessing aggregates, hiding persistence details from the domain layer.
**Interface** (defined in the domain layer):
```
interface OrderRepository {
findById(id: OrderId): Order | null
save(order: Order): void
findByCustomer(customerId: CustomerId): Order[]
}
```
**Rules**:
- One repository per aggregate root (not per entity or table)
- Repository interface lives in the domain layer; implementation lives in infrastructure
- Repositories return fully constituted aggregates, not partial data
- Do not put query logic in repositories — use separate read models for complex queries
## Event Storming
A collaborative workshop technique for discovering domain events, commands, and aggregates.
### Process
1. **Gather**: Domain experts + developers in a room with unlimited sticky notes
2. **Domain Events** (orange): Brainstorm everything that happens in the domain (past tense)
3. **Commands** (blue): What triggers each event? (imperative: "Place Order")
4. **Actors** (yellow): Who or what issues the command? (User, System, Scheduler)
5. **Aggregates** (pale yellow): Group events around the entity they affect
6. **Bounded Contexts**: Draw boundaries around related aggregate clusters
7. **Policies** (purple): Automated reactions ("When OrderPlaced, then ReserveInventory")
### Output
A visual map of the domain that reveals:
- Core business processes and their interactions
- Natural bounded context boundaries
- Integration points between contexts
- Hot spots (areas of complexity, contention, or confusion)
## Design Heuristics
- If two concepts change for different reasons, they belong in different bounded contexts
- If you need a transaction across two aggregates, reconsider the aggregate boundaries
- Start with larger aggregates and split when you encounter contention or performance issues
- Domain events are the primary integration mechanism between bounded contexts
- Ubiquitous language: use the same terms in code, documentation, and conversation with domain experts
@@ -0,0 +1,92 @@
# NFR Design Guide
## Resilience Patterns
Apply these patterns based on failure mode analysis:
| Pattern | When to Use | Key Configuration |
|---------|-------------|-------------------|
| **Circuit Breaker** | External service calls, database connections | Failure threshold (5), timeout (30s), half-open retry interval (60s) |
| **Bulkhead** | Isolating failures between subsystems | Thread pool per dependency, max concurrent requests, queue depth |
| **Retry with Backoff** | Transient failures (network blips, rate limits) | Max retries (3), exponential backoff (100ms, 200ms, 400ms), jitter |
| **Timeout** | Every external call without exception | Connect timeout (5s), read timeout (30s), total timeout (60s) |
| **Fallback** | Degraded mode is acceptable over total failure | Static fallback, cached response, default value, feature flag |
| **Rate Limiter** | Protecting downstream from upstream bursts | Token bucket or sliding window, per-user and global limits |
### Circuit Breaker State Machine
- **Closed** (normal): Requests pass through. Count failures. Open when threshold reached.
- **Open** (tripped): All requests fail fast without calling downstream. Wait for reset timeout.
- **Half-Open** (probing): Allow one request through. If it succeeds, close. If it fails, re-open.
## Caching Architecture
### Cache Placement Decision Matrix
| Scenario | Cache Location | TTL Strategy | Invalidation |
|----------|---------------|--------------|--------------|
| Static assets | CDN edge | Long (24h+) | Version hash in URL |
| API responses (read-heavy) | Reverse proxy / API gateway | Medium (5-15 min) | Event-driven purge |
| Database query results | Application-level (Redis/Memcached) | Short (1-5 min) | Write-through or write-behind |
| Session data | Distributed cache | Session lifetime | Explicit delete on logout |
| Computed aggregations | Materialized view / pre-computed cache | Scheduled refresh | Rebuild on source change |
### Cache Consistency Rules
- Never cache data that must be strongly consistent across requests
- Use cache-aside (lazy loading) as the default pattern
- Write-through only when write latency tolerance allows it
- Always set a TTL -- unbounded caches become stale data stores
- Monitor cache hit ratio (target >90% for read-heavy paths)
## Scalability Patterns
### Horizontal Scaling Strategies
| Strategy | Best For | Considerations |
|----------|----------|----------------|
| **Stateless services** | API servers, web frontends | No session affinity needed; scale by adding instances behind load balancer |
| **Sharding** | Large datasets, multi-tenant systems | Choose shard key carefully (tenant ID, region); avoid cross-shard queries |
| **Read replicas** | Read-heavy workloads (>10:1 read:write) | Accept replication lag; route writes to primary, reads to replicas |
| **Event-driven decoupling** | Bursty workloads, async processing | Use message queues; consumers scale independently of producers |
| **CQRS** | Different read/write scaling needs | Separate read and write models; accept eventual consistency |
### Shard Key Selection Criteria
- High cardinality (many distinct values)
- Even distribution (no hot partitions)
- Query locality (most queries target a single shard)
- Immutable (changing shard keys requires data migration)
## Reliability Engineering
### SLI/SLO Definition Template
For each critical user journey, define:
- **SLI** (Service Level Indicator): The metric being measured
- Availability: `successful_requests / total_requests`
- Latency: `requests_below_threshold / total_requests` (e.g., p99 < 500ms)
- Correctness: `correct_responses / total_responses`
- **SLO** (Service Level Objective): The target for the SLI
- Format: "99.9% of requests return successfully over a 30-day window"
- **Error Budget**: `1 - SLO` = allowable failure rate
- 99.9% SLO = 43.2 minutes of downtime per 30 days
### Failure Mode Checklist
For each component, assess:
- [ ] What happens when this component is unavailable?
- [ ] What happens when response time doubles?
- [ ] What happens when throughput exceeds capacity?
- [ ] What happens when a dependency returns corrupted data?
- [ ] Is there a graceful degradation path?
- [ ] What is the blast radius of a failure? (single user, tenant, region, global)
- [ ] What is the recovery procedure? (automatic, manual, requires restart)
## NFR-to-Architecture Mapping
| NFR Category | Architectural Implication |
|-------------|--------------------------|
| Latency < 100ms | In-memory cache, CDN, connection pooling, async non-blocking I/O |
| Availability > 99.9% | Multi-AZ deployment, health checks, auto-restart, circuit breakers |
| Throughput > 10K rps | Horizontal scaling, load balancing, connection pooling, async processing |
| Data durability | Replicated storage, point-in-time recovery, write-ahead logging |
| Disaster recovery RTO < 1h | Multi-region active-passive, automated failover, tested runbooks |
| Zero-downtime deployments | Blue-green or canary deploys, backward-compatible migrations, feature flags |
@@ -0,0 +1,149 @@
# Non-Functional Requirement Design Patterns
## Purpose
Patterns for achieving reliability, performance, and resilience in production systems. These patterns address the gap between "it works" and "it works reliably at scale."
## Caching Strategies
### Cache-Aside (Lazy Loading)
```
Read: Check cache -> if miss, read from DB -> populate cache -> return
Write: Write to DB -> invalidate cache
```
- Most common pattern; application manages cache explicitly
- Risk: Cache miss thundering herd under cold start or cache failure
- Mitigation: Use cache warming on startup and request coalescing
### Write-Through
```
Write: Write to cache -> cache synchronously writes to DB
Read: Always read from cache
```
- Cache is always consistent with DB
- Higher write latency (two writes on every mutation)
- Best when reads vastly outnumber writes
### Write-Behind (Write-Back)
```
Write: Write to cache -> cache asynchronously writes to DB (batched)
Read: Always read from cache
```
- Lowest write latency (only cache write is synchronous)
- Risk: Data loss if cache fails before async write completes
- Best for high-write throughput where brief inconsistency is acceptable
### Cache Invalidation Rules
- Set TTL (time-to-live) on every cached item — stale data is worse than a cache miss
- Use explicit invalidation on writes when consistency matters
- Never cache error responses (negative caching requires very short TTLs)
- Monitor cache hit ratio — below 80% indicates the cache is not helping
## Circuit Breaker
### Purpose
Prevent cascading failures when a downstream service is unavailable.
### States
1. **Closed** (normal): Requests flow through. Track failure count.
2. **Open** (tripped): All requests fail immediately without calling downstream. Return fallback or error.
3. **Half-Open** (testing): Allow a limited number of probe requests. If they succeed, return to Closed. If they fail, return to Open.
### Configuration
- **Failure threshold**: Number of consecutive failures before opening (e.g., 5)
- **Open duration**: Time to wait before transitioning to half-open (e.g., 30 seconds)
- **Probe count**: Number of test requests in half-open state (e.g., 3)
### Implementation Notes
- Circuit breakers should be per-dependency, not global
- Log state transitions for observability
- Provide meaningful fallback behavior (cached data, degraded response, queue for retry)
## Bulkhead
### Purpose
Isolate failures to prevent one failing component from consuming all system resources.
### Approaches
- **Thread pool isolation**: Each dependency gets its own thread pool with a fixed size. If dependency A exhausts its pool, dependency B is unaffected.
- **Connection pool isolation**: Separate connection pools per downstream service.
- **Process isolation**: Run critical and non-critical workloads in separate processes or containers.
### Sizing Rule
Size each bulkhead based on the dependency's expected throughput plus a buffer. Too small causes unnecessary rejection; too large defeats the purpose.
## Retry with Exponential Backoff
### Pattern
```
Retry after: base_delay * 2^attempt + random_jitter
Example: 100ms, 200ms, 400ms, 800ms, 1600ms (+ jitter)
```
### Rules
- **Always add jitter** — without it, retries from multiple clients synchronize and create thundering herd
- **Set a maximum retry count** (typically 3-5) — infinite retries cause resource exhaustion
- **Only retry transient failures** — do not retry 400 Bad Request or 403 Forbidden
- **Retryable errors**: 429 (rate limited), 500, 502, 503, 504, connection timeout, connection reset
- **Ensure idempotency** — retried operations must produce the same result (use idempotency keys)
## Rate Limiting
### Algorithms
- **Token Bucket**: Accumulate tokens over time; each request consumes a token. Allows controlled bursts.
- **Sliding Window**: Count requests in a rolling time window. Smoother than fixed windows (avoids boundary bursts).
- **Fixed Window**: Count requests per time interval. Simplest but allows 2x burst at window boundaries.
### Application
- Apply at API gateway level for external consumers
- Apply per-user or per-tenant for fair usage
- Return HTTP 429 with `Retry-After` header
- Communicate rate limits in response headers (`X-RateLimit-Remaining`, `X-RateLimit-Reset`)
## Load Balancing
### Algorithms
- **Round Robin**: Distribute evenly across instances. Simple, works when instances are homogeneous.
- **Least Connections**: Route to the instance with fewest active connections. Better when requests have variable duration.
- **Weighted**: Assign weights based on instance capacity. Use when instances have different sizes.
- **Consistent Hashing**: Route based on request key (user ID, session). Maintains affinity for caching benefits.
### Health Checks
- **Shallow**: TCP or HTTP 200 check (is the process running?)
- **Deep**: Check downstream dependencies (can the service actually serve requests?)
- Use shallow checks for load balancer routing; deep checks for alerting
## Connection Pooling
### Purpose
Reuse expensive connections (database, HTTP, gRPC) instead of creating new ones per request.
### Configuration
- **Minimum pool size**: Connections kept warm during idle periods (e.g., 5)
- **Maximum pool size**: Upper bound to prevent resource exhaustion (e.g., 20)
- **Connection timeout**: How long to wait for a connection from the pool (e.g., 5 seconds)
- **Idle timeout**: Close connections unused for this duration (e.g., 10 minutes)
- **Max lifetime**: Close connections after this age regardless of use (e.g., 30 minutes) to prevent stale connections
### Sizing Formula
```
Pool size = (Requests per second) x (Average request duration in seconds) x 1.5 (buffer)
```
## Graceful Degradation
### Principle
When a dependency fails, reduce functionality rather than failing entirely.
### Strategies
- **Feature flags**: Disable non-critical features when their backing service is down
- **Cached fallback**: Serve stale cached data with a "data may be outdated" indicator
- **Default values**: Use sensible defaults when personalization service is unavailable
- **Queue for later**: Accept writes into a queue when the write path is degraded
- **Read-only mode**: Disable writes but keep reads functioning
### Priority Tiers
1. **Critical path**: Must always work (authentication, core transaction)
2. **Important**: Degrade gracefully (recommendations, analytics, notifications)
3. **Nice to have**: Disable entirely under pressure (social features, cosmetic enhancements)
Map every dependency to a tier and define the degradation strategy per tier before production launch.
@@ -0,0 +1,118 @@
# Reviewing Artifacts (Architecture Lens)
When invoked as a reviewer, your role changes. You are NOT designing — you are evaluating someone else's design with fresh eyes.
## Stance
- You did not produce this work. Judge the output independently.
- Your scope is the artifacts you were passed plus the shared contracts named in the invocation prompt - the current unit and its declared upstream, not the whole project's history. Cross-unit contract verification runs against those shared contracts, not by reading other units' design directories.
- You do not have access to the builder's reasoning (plan.md, memory.md). This is intentional.
- Your job is to find architectural unsoundness, broken cross-references, missing concerns, and designs that won't survive implementation.
- "READY" means a developer could implement from this without guessing. Not perfect — implementable.
## What to Check
### Application/Domain Design
- Component boundaries clear? (what owns what?)
- Dependencies correct and complete? (hidden couplings?)
- Circular dependencies?
- Single responsibility per component? (no god-components)
- Entity relationships correct? (cardinality, direction)
### Functional Design
- All business rules complete? (trigger, logic, violation for each)
- Entities have all attributes needed to implement rules?
- State machines complete? (all states reachable, no dead ends)
- API specs cover error cases, not just happy paths?
- Cross-unit contract boundaries respected? Verify against the shared inception contracts passed with the invocation (`components.md`, `contract-summary.md`, `unit-of-work.md`), NOT against sibling units' `construction/<other-unit>/functional-design/` prose and not via grep, glob, or shell patterns that span sibling unit paths. If the current unit's design names a specific integration point in another unit, open the owning file (resolved via the shared contracts, not by browsing or searching the sibling unit's directory) to spot-check; do not sweep the sibling unit.
### NFR Design
- Quality targets measurable? (SLOs with numbers)
- Technology choices justified against NFRs?
- Alternatives documented with trade-off reasoning?
- Cost model realistic at scale?
- Security boundaries defined?
### Infrastructure Design
- Every component mapped to infrastructure?
- Networking complete? (ingress, egress, inter-service)
- DR strategy with RTO/RPO?
- Scaling triggers and limits defined?
- Cost estimate present?
### Units Generation
- Unit boundaries clean? (minimal cross-unit deps)
- Dependency graph acyclic?
- Stories mapped completely? (no orphans)
- Each unit independently deployable?
### Validation Tools
If the stage definition lists validation tools, **run them via shell** before writing your review. Include results in findings. Interpret them — a tool failure might be acceptable with documented rationale.
## How to Lodge Review Comments
Write your review to the review file the dispatch names (the `reviewFile` path
the request returned, under the intent record's `.aidlc-reviews/` directory).
That file is the only thing you write: never edit the artifact you are
reviewing or any other stage output. The engine records your review beside the
artifact and refuses a verdict whose artifacts changed. `ID` values are
stable (`R-01`, `R-02`, ...): never renumber, reuse, or change an existing ID.
`Location` MUST be a workspace-relative artifact path followed by the exact
section or element. `Required action` MUST state the concrete work in plain
language. On the first review, every finding has status `New`.
Use this exact format:
```markdown
## Review
**Verdict:** READY | NOT-READY
**Reviewer:** aidlc-architecture-reviewer-agent
**Date:** [ISO timestamp from Bash]
**Iteration:** [1, 2, etc.]
### Findings
| ID | Severity | Location | Finding | Required action | Status |
|---|---|---|---|---|---|
| R-01 | Critical | aidlc/spaces/<space>/intents/<intent-record>/inception/domain-design/components.md > component CMP-003 dependencies | CMP-003 depends on CMP-001 which depends on CMP-003, creating a cycle | Break the cycle, for example by extracting the shared concern into a new component | New |
| R-02 | Major | aidlc/spaces/<space>/intents/<intent-record>/construction/<unit>/functional-design/entities.md > entity ENT-005 | ENT-005 references entity "Payment", which is not defined | Define Payment in the owning artifact or reference the correct upstream entity | New |
| R-03 | Minor | aidlc/spaces/<space>/intents/<intent-record>/construction/<unit>/nfr-design/performance-design.md > Caching layer cost | No cost estimate exists for the caching layer | Add a cost estimate or explicitly record it as TBD with an owner | New |
### Validation Tool Results
| Tool | Result | Interpretation |
|---|---|---|
| validate-domain-model | FAIL: circular dep CMP-003↔CMP-001 | Confirms finding R-01 — must fix |
| validate-entities | PASS | All IDs unique, refs valid |
### Summary
[1-2 sentences: what's the main architectural concern, or why it's ready.]
```
For the `Date` field, obtain a real UTC timestamp by running `date -u +"%Y-%m-%dT%H:%M:%SZ"` in the shell and paste the actual output. Never guess or infer the date.
### Severity Levels
| Severity | Meaning | Blocks READY? |
|---|---|---|
| Critical | Architectural flaw that will cause failure at implementation or runtime | Yes |
| Major | Design gap that will cause significant rework | Yes (if >2 major) |
| Minor | Could be better, not blocking | No |
### Verdict Rules
- **READY** if: zero Critical, ≤2 Major, any number of Minor
- **NOT-READY** if: any Critical, OR >2 Major findings
### On Subsequent Iterations
When the dispatch brief includes `Prior findings (carry IDs forward)`:
- Treat that table as authoritative for prior human dispositions; it is
rendered from the audit ledger without rewriting the reviewed artifact.
- Reproduce every prior row with the same ID; never renumber, reuse, or drop an ID.
- Re-check the cited location and set `Status` to exactly one of `Unresolved`, `Resolved`, `Rejected: <reason>`, or `Accepted risk`. A partial fix remains `Unresolved`, with `Required action` narrowed to the work still needed.
- Preserve a `Rejected: <reason>` or `Accepted risk` disposition only when the prior-findings input carries it; do not invent either disposition.
- Add a genuinely new finding only under the next unused `R-NN` ID and mark it `New`.
- Write the whole review afresh to the review file named for this iteration; it carries every prior row plus any new ones, never a second table.
@@ -0,0 +1,195 @@
# AWS CDK Best Practices
## Purpose
Guidelines for building maintainable, secure, and testable infrastructure using the AWS Cloud Development Kit. These practices apply to CDK v2 with TypeScript (the recommended language for most teams).
## Construct Levels
### L1 (Cfn Resources)
- Direct CloudFormation resource wrappers (e.g., `CfnBucket`)
- Use only when L2 constructs do not expose a needed property
- Require manual configuration of all properties (no defaults)
### L2 (Curated Constructs)
- AWS-maintained constructs with sensible defaults (e.g., `Bucket`, `Function`, `Table`)
- Include helper methods (e.g., `bucket.grantRead(lambda)`)
- Preferred for most use cases — they encode AWS best practices
### L3 (Patterns)
- Higher-level constructs combining multiple resources (e.g., `LambdaRestApi`)
- Use when the pattern fits your needs exactly
- Avoid if you need significant customization — drop down to L2 instead
## Construct Design Patterns
### Single Responsibility
Each custom construct should represent one logical unit (a service, a data pipeline stage, a monitoring stack). Do not create constructs that build unrelated resources.
### Props Interface Pattern
```typescript
export interface OrderServiceProps {
readonly vpc: ec2.IVpc;
readonly table: dynamodb.ITable;
readonly environment: string; // 'dev' | 'staging' | 'prod'
readonly alarmTopic?: sns.ITopic; // optional props use ?
}
export class OrderService extends Construct {
public readonly api: apigateway.RestApi; // expose outputs as public readonly
constructor(scope: Construct, id: string, props: OrderServiceProps) {
super(scope, id);
// ...
}
}
```
### Rules
- Accept dependencies via props (dependency injection), do not create shared resources inside constructs
- Use interface types (`IVpc`, `ITable`) for props, not concrete types — enables cross-stack references
- Expose outputs as public readonly properties for consuming constructs
- Prefix optional props with documentation explaining the default behavior
## Stack Organization
### Recommended Structure
```
/infrastructure
/bin
app.ts # CDK app entry point, environment configuration
/lib
/constructs # Reusable L3 constructs
order-service.ts
monitoring.ts
/stacks
network-stack.ts # VPC, subnets, security groups
data-stack.ts # DynamoDB, S3, RDS
compute-stack.ts # Lambda, ECS, API Gateway
monitoring-stack.ts # CloudWatch, alarms, dashboards
/test
order-service.test.ts
data-stack.test.ts
```
### Stack Separation Guidelines
- Separate stacks by lifecycle: resources that change together should be in the same stack
- Stateful resources (databases, S3 buckets) in separate stacks from stateless (Lambda, API Gateway)
- Stateful stacks change rarely; stateless stacks deploy frequently
- Use cross-stack references sparingly — they create deployment coupling
## Environment-Aware Stacks
### Pattern
```typescript
// bin/app.ts
const app = new cdk.App();
const env = app.node.tryGetContext('env') || 'dev';
const config = {
dev: { instanceType: 't3.small', minCapacity: 1, maxCapacity: 2 },
staging: { instanceType: 't3.medium', minCapacity: 2, maxCapacity: 4 },
prod: { instanceType: 't3.large', minCapacity: 3, maxCapacity: 10 },
}[env];
new ComputeStack(app, `ComputeStack-${env}`, {
env: { account: process.env.CDK_DEFAULT_ACCOUNT, region: 'us-east-1' },
config,
});
```
### Rules
- Never hardcode account IDs or regions — use environment variables or context
- Use the same code for all environments; parameterize differences through config
- Production stacks must specify explicit `env` (account + region) — do not rely on defaults
## CDK Testing
### Assertion Tests (Fine-Grained)
```typescript
test('DynamoDB table has encryption enabled', () => {
const app = new cdk.App();
const stack = new DataStack(app, 'TestStack');
const template = Template.fromStack(stack);
template.hasResourceProperties('AWS::DynamoDB::Table', {
SSESpecification: {
SSEEnabled: true,
},
});
});
```
### Snapshot Tests (Regression Detection)
```typescript
test('stack matches snapshot', () => {
const app = new cdk.App();
const stack = new DataStack(app, 'TestStack');
const template = Template.fromStack(stack);
expect(template.toJSON()).toMatchSnapshot();
});
```
- Update snapshots intentionally (`jest --updateSnapshot`) after deliberate changes
- Review snapshot diffs in pull requests — they show exactly what infrastructure changes
### What to Test
- Security properties: encryption enabled, public access blocked, least-privilege policies
- Critical configuration: retention policies, backup settings, auto-scaling parameters
- Resource counts: expected number of Lambda functions, tables, queues
- Do NOT test CDK internals or CloudFormation implementation details
## Security Defaults
### Encryption
- S3: `encryption: s3.BucketEncryption.S3_MANAGED` (minimum) or KMS for sensitive data
- DynamoDB: `encryption: dynamodb.TableEncryption.AWS_MANAGED` (default) or customer-managed KMS
- SQS: `encryption: sqs.QueueEncryption.KMS` for sensitive message content
- EBS: Enable encryption by default in account settings
### Access Control
- S3: `blockPublicAccess: s3.BlockPublicAccess.BLOCK_ALL` (always, unless serving public static content)
- Lambda: Use `grant*` methods instead of writing IAM policies manually
- API Gateway: Add authorization on every route (IAM, Cognito, or Lambda authorizer)
### Logging
- S3: Enable access logging to a dedicated logging bucket
- API Gateway: Enable access logging and execution logging
- Lambda: Logs go to CloudWatch automatically; set retention (`logRetention: logs.RetentionDays.ONE_MONTH`)
### Least Privilege
```typescript
// Good: specific grant
table.grantReadData(lambdaFunction);
// Bad: overly broad
lambdaFunction.addToRolePolicy(new iam.PolicyStatement({
actions: ['dynamodb:*'],
resources: ['*'],
}));
```
## CDK Aspects for Compliance
### Purpose
Aspects visit every construct in the tree and can validate, warn, or modify resources.
```typescript
class EncryptionChecker implements cdk.IAspect {
public visit(node: IConstruct): void {
if (node instanceof s3.CfnBucket) {
if (!node.bucketEncryption) {
Annotations.of(node).addError('S3 bucket must have encryption enabled');
}
}
}
}
// Apply to the entire app
Aspects.of(app).add(new EncryptionChecker());
```
### Common Compliance Aspects
- Verify all S3 buckets have encryption and block public access
- Verify all DynamoDB tables have point-in-time recovery enabled
- Verify all Lambda functions have reserved concurrency set
- Verify all security groups do not allow 0.0.0.0/0 ingress
- Tag all resources with required cost allocation tags
@@ -0,0 +1,142 @@
# AWS Cost Optimization Patterns
## Purpose
Practical strategies for reducing AWS costs without sacrificing reliability or performance. Cost optimization is an ongoing discipline, not a one-time exercise.
## Compute: Rightsizing
### Process
1. Enable AWS Compute Optimizer (free, account-wide)
2. Review recommendations after 14 days of data collection
3. Identify over-provisioned instances (CPU < 20%, memory < 30% on average)
4. Resize in non-production first, then production during maintenance windows
### Common Findings
- Most teams over-provision by 30-50% at initial deployment
- Graviton (ARM) instances offer 20-40% better price-performance than x86 equivalents
- Consider burstable instances (t3/t4g) for workloads with variable CPU patterns
### Action Items
- Review instance utilization monthly via Cost Explorer or Compute Optimizer
- Set CloudWatch alarms for sustained low utilization (< 10% CPU for 7 days)
- Automate non-production instance scheduling (stop at 7 PM, start at 7 AM)
## Pricing Models
### On-Demand
- No commitment, highest per-hour cost
- Use for: unpredictable workloads, short-term spikes, new applications before usage patterns are established
### Savings Plans
- 1-year or 3-year commitment to a consistent amount of compute usage (measured in $/hour)
- **Compute Savings Plans**: Apply across EC2, Lambda, and Fargate (most flexible)
- **EC2 Instance Savings Plans**: Locked to instance family and region (deeper discount)
- Typical savings: 30-40% (1-year, no upfront) to 60-72% (3-year, all upfront)
- Start with Compute Savings Plans covering your baseline usage; use On-Demand for peaks
### Reserved Instances
- 1-year or 3-year commitment to specific instance type, region, and OS
- Less flexible than Savings Plans; similar discounts
- Consider only for steady-state workloads with very predictable instance types
### Spot Instances
- Up to 90% discount, but can be interrupted with 2-minute notice
- Use for: batch processing, CI/CD workers, data processing, stateless web servers behind auto-scaling groups
- Best practice: Diversify across multiple instance types and AZs to reduce interruption frequency
- Never use for: databases, single-instance workloads, or anything that cannot tolerate interruption
## Storage: S3 Lifecycle Policies
### Recommended Transitions
```
S3 Standard (active data, frequent access)
→ 30 days → S3 Intelligent-Tiering (variable access patterns)
→ 90 days → S3 Infrequent Access (known infrequent access)
→ 180 days → S3 Glacier Instant Retrieval (rare access, millisecond retrieval)
→ 365 days → S3 Glacier Deep Archive (archive, 12-hour retrieval)
```
### Rules
- Analyze access patterns with S3 Storage Lens before setting lifecycle rules
- Use S3 Intelligent-Tiering for unpredictable access patterns (automates transitions)
- Delete incomplete multipart uploads after 7 days (they accumulate silently)
- Enable S3 analytics to validate lifecycle policy effectiveness
### Quick Wins
- Delete old CloudTrail logs in S3 after compliance retention period
- Transition ELB/CloudFront access logs to Glacier after 90 days
- Compress objects before upload (gzip, zstd) — reduces storage and transfer costs
## Database: DynamoDB
### On-Demand vs Provisioned Capacity
| Factor | On-Demand | Provisioned |
|--------|-----------|-------------|
| Traffic pattern | Unpredictable, spiky | Steady, predictable |
| Pricing | Per-request | Per-capacity-unit-hour |
| Scaling | Instant (within limits) | Auto-scaling with lag |
| Best for | New tables, dev/test, event-driven | Production with known patterns |
### Cost Reduction Strategies
- Use provisioned capacity with auto-scaling for steady-state production tables
- Enable reserved capacity for predictable base load (additional discount on provisioned)
- Use TTL to automatically delete expired items (no write cost for TTL deletions)
- Design partition keys to distribute load evenly (hot partitions waste provisioned capacity)
## Lambda Cost Optimization
### Memory and Duration
- Lambda pricing = (memory allocated) x (execution duration) x (number of invocations)
- More memory also means more CPU — increasing memory can reduce duration and total cost
- Use AWS Lambda Power Tuning to find the optimal memory setting per function
### Strategies
- Minimize cold starts: keep deployment packages small, use layers for shared dependencies
- Use ARM/Graviton runtime (`arm64`) for 20% cost reduction and better performance
- Set appropriate timeout (not maximum 15 minutes for a function that runs in 3 seconds)
- Batch process SQS messages (receive up to 10 messages per invocation)
- Avoid provisioned concurrency unless latency requirements demand it (it is expensive)
### Invocation Reduction
- Use SQS batch window to accumulate messages before invoking Lambda
- Use EventBridge rules with content filtering to invoke only for relevant events
- Cache results in DynamoDB or ElastiCache to reduce redundant compute
## Cost Allocation Tagging
### Required Tags
Define and enforce a minimum tag set across all resources:
```
Project: project-name
Environment: dev | staging | prod
Team: team-name
CostCenter: cost-center-code
Service: service-name
```
### Enforcement
- Use AWS Organizations SCPs to deny resource creation without required tags
- Use CDK Aspects to add tags automatically and validate tag presence
- Enable cost allocation tags in Billing console (tags must be activated to appear in Cost Explorer)
## Cost Explorer Queries
### Monthly Cost Reviews
1. **Cost by service**: Identify top 5 cost drivers
2. **Cost by tag (team)**: Attribute costs to responsible teams
3. **Daily cost trend**: Detect unexpected cost spikes
4. **Cost by usage type**: Identify specific resources (data transfer, API calls, storage)
### Anomaly Detection
- Enable AWS Cost Anomaly Detection for automatic notification of unexpected cost increases
- Set budget alerts at 50%, 80%, and 100% of expected monthly spend
- Create separate budgets per environment (production vs non-production)
### Monthly Cost Optimization Checklist
- [ ] Review Compute Optimizer recommendations for rightsizing
- [ ] Check for idle resources (unused EIPs, unattached EBS volumes, idle load balancers)
- [ ] Review Savings Plans utilization and coverage
- [ ] Check S3 storage distribution across tiers
- [ ] Review data transfer costs (cross-region, internet egress)
- [ ] Verify non-production environments are scheduled for off-hours shutdown
- [ ] Review Lambda function memory settings with Power Tuning results
@@ -0,0 +1,108 @@
# Infrastructure Guide
## IaC Tool Selection
| Tool | Best For | Considerations |
|------|----------|---------------|
| AWS CDK | AWS-native, TypeScript/Python teams, complex constructs | Vendor lock-in, steep learning curve |
| Terraform | Multi-cloud, team standardization, mature ecosystem | HCL syntax, state management complexity |
| CloudFormation | AWS-only, simple stacks, when CDK is overkill | Verbose YAML/JSON, slow drift detection |
| Pulumi | Polyglot teams wanting general-purpose languages | Smaller community, state backend choice |
| Docker Compose | Local development, simple multi-container apps | Not for production orchestration |
## CI/CD Pipeline Design
### Standard Pipeline Stages
```
[Source] -> [Lint] -> [Build] -> [Unit Test] -> [SAST] -> [Package] ->
[Integration Test] -> [Deploy Staging] -> [E2E Test] -> [Security Scan] ->
[Approval Gate] -> [Deploy Production] -> [Smoke Test] -> [Monitor]
```
### Stage Requirements
- **Lint**: Fail fast on formatting and static analysis violations. Under 30 seconds.
- **Build**: Compile/transpile, resolve dependencies. Cache aggressively. Under 2 minutes.
- **Unit Test**: Full suite. Fail the pipeline on any failure. Under 3 minutes.
- **SAST**: Static security scan. Block on high/critical findings. Under 5 minutes.
- **Package**: Build container image or deployment artifact. Tag with commit SHA.
- **Integration Test**: Run against test database and mock external services. Under 10 minutes.
- **Deploy Staging**: Automated. Mirror production topology at reduced scale.
- **E2E Test**: Critical paths only. Under 15 minutes. Flaky tests quarantined, not skipped.
- **Security Scan**: DAST against staging. Dependency vulnerability check.
- **Approval Gate**: Manual approval for production (optional, based on risk appetite).
- **Deploy Production**: Automated with selected deployment strategy.
- **Smoke Test**: Verify core endpoints respond correctly post-deploy. Under 2 minutes.
## Deployment Strategies
### Blue-Green
- Two identical environments; traffic switches atomically
- **Pro**: Instant rollback, zero downtime
- **Con**: Double infrastructure cost, database schema sync complexity
- **Use when**: Zero-downtime required, database changes are backward-compatible
### Canary
- Route small percentage of traffic (1-5%) to new version, gradually increase
- **Pro**: Limited blast radius, real-user validation
- **Con**: Requires traffic splitting, metric comparison automation
- **Use when**: High-traffic systems, need to validate under real load
### Rolling
- Replace instances incrementally (1 at a time or N at a time)
- **Pro**: No extra infrastructure, gradual rollout
- **Con**: Mixed versions running simultaneously, slower rollback
- **Use when**: Stateless services, backward-compatible changes
### Recreate
- Stop all old instances, start all new instances
- **Pro**: Simple, no version mixing
- **Con**: Downtime during transition
- **Use when**: Acceptable maintenance window, breaking changes
## Monitoring & Observability Stack
### The Four Pillars
1. **Metrics**: Numeric measurements over time (CPU, latency, error count)
- Tool examples: CloudWatch, Prometheus + Grafana, Datadog
- Key metrics: RED (Rate, Errors, Duration) for services; USE (Utilization, Saturation, Errors) for resources
2. **Logs**: Structured event records from application and infrastructure
- Format: JSON with timestamp, level, service, traceId, message, context
- Tool examples: CloudWatch Logs, ELK stack, Loki
- Retention: 30 days hot, 90 days warm, 1 year cold
3. **Traces**: Request flow across services
- Tool examples: X-Ray, Jaeger, Zipkin, OpenTelemetry
- Instrument: HTTP handlers, database calls, external API calls, queue operations
4. **Alerts**: Automated notifications for anomalies
- Alert on symptoms (error rate > 1%), not causes (CPU > 80%)
- Severity levels: P1 (page), P2 (ticket), P3 (dashboard)
- Include runbook link in every alert
## Container Orchestration Checklist
For containerized deployments, define:
- [ ] Base image selection (minimal, security-patched, pinned version)
- [ ] Multi-stage build for smaller production images
- [ ] Health check endpoint (`/health` or `/readyz`)
- [ ] Graceful shutdown handling (SIGTERM, drain connections)
- [ ] Resource limits (CPU, memory) to prevent noisy-neighbor issues
- [ ] Secrets management (not in image, not in env vars -- use secrets manager)
- [ ] Log output to stdout/stderr (not file-based)
- [ ] Non-root user in container
- [ ] Read-only filesystem where possible
## Environment Strategy
```
Local Dev -> Developer laptop, docker-compose, hot reload
CI -> Ephemeral, created per pipeline run, destroyed after
Staging -> Persistent, mirrors production topology, reduced scale
Production -> Full scale, multi-AZ, monitoring and alerting active
```
Parity rules:
- Staging MUST use the same IaC templates as production (parameterized for scale)
- Staging MUST use the same database engine and version as production
- Staging SHOULD have representative (anonymized) data volume
@@ -0,0 +1,145 @@
# AWS Well-Architected Framework
## Purpose
A structured approach for evaluating architectures against AWS best practices across six pillars. Use this framework during architecture reviews, before production launches, and periodically for existing workloads.
## Six Pillars Overview
| Pillar | Focus | Key Metric |
|--------|-------|------------|
| Operational Excellence | Run and monitor systems effectively | Mean time to recovery (MTTR) |
| Security | Protect data, systems, and assets | Number of security findings |
| Reliability | Recover from failures, meet demand | Availability percentage (e.g., 99.9%) |
| Performance Efficiency | Use resources efficiently | Latency percentiles (p50, p95, p99) |
| Cost Optimization | Avoid unnecessary costs | Cost per transaction/user |
| Sustainability | Minimize environmental impact | Resources per unit of work |
## Pillar 1: Operational Excellence
### Key Questions
- How do you determine what your priorities are?
- How do you design your workload to understand its state?
- How do you reduce defects, ease remediation, and improve flow?
- How do you evolve your operations?
### Best Practices
- **Infrastructure as code**: All resources defined in CDK/CloudFormation, version controlled
- **Observability**: Structured logging, distributed tracing (X-Ray), custom metrics (CloudWatch)
- **Runbooks and playbooks**: Documented procedures for common operational tasks and incident response
- **Deployment automation**: CI/CD pipelines with automated testing, canary deployments, automatic rollback
- **Game days**: Regularly simulate failures to validate operational readiness
### Common Anti-Patterns
- Manual infrastructure changes via the console
- Logging only errors (missing context for debugging)
- No runbooks for common failure modes
- Deploying on Friday afternoons without monitoring
## Pillar 2: Security
### Key Questions
- How do you manage identities and permissions?
- How do you detect and investigate security events?
- How do you protect your network, compute, and data?
### Best Practices
- **Identity**: Use IAM roles (not long-lived keys), enforce MFA, apply least-privilege policies
- **Detection**: Enable CloudTrail, GuardDuty, Security Hub, and Config rules
- **Data protection**: Encrypt at rest (KMS) and in transit (TLS 1.2+), classify data sensitivity
- **Network**: Use VPC with private subnets for databases/compute, security groups as firewalls, VPC endpoints for AWS services
- **Incident response**: Pre-provisioned forensic tools, automated containment playbooks
### Common Anti-Patterns
- Wildcard IAM policies (`Action: "*"`, `Resource: "*"`)
- Secrets in environment variables or code (use Secrets Manager or Parameter Store)
- Public S3 buckets or security groups open to 0.0.0.0/0
- No encryption on databases or message queues
## Pillar 3: Reliability
### Key Questions
- How do you manage service quotas and constraints?
- How does your workload adapt to changes in demand?
- How do you design interactions to prevent failures?
- How do you test reliability?
### Best Practices
- **Multi-AZ**: Deploy across at least 2 Availability Zones for all critical components
- **Auto scaling**: Configure for both scale-out and scale-in with appropriate cooldowns
- **Fault isolation**: Use bulkheads, circuit breakers, and timeouts for all remote calls
- **Backup and recovery**: Automated backups with tested restore procedures, define RPO/RTO
- **Chaos engineering**: Inject failures (instance termination, AZ loss, latency) to verify resilience
### Common Anti-Patterns
- Single-AZ deployments for production workloads
- No health checks on load balancer targets
- Untested backup restores (backups exist but recovery has never been validated)
- Hard dependencies on services without fallback behavior
## Pillar 4: Performance Efficiency
### Key Questions
- How do you select the best performing architecture?
- How do you select and manage your compute, storage, and database solutions?
- How do you monitor to ensure performance?
### Best Practices
- **Right-size resources**: Start small, measure, and adjust — do not guess capacity
- **Caching**: CloudFront for static content, ElastiCache for application data, API Gateway caching
- **Database selection**: Match engine to access pattern (relational for joins, DynamoDB for key-value, OpenSearch for full-text)
- **Async processing**: Offload long-running tasks to SQS/Lambda, keep API response times fast
- **Load testing**: Establish baseline performance and test at 2-3x expected peak load
### Common Anti-Patterns
- Using RDS for simple key-value lookups (DynamoDB is more efficient)
- Over-provisioned instances running at 5% utilization
- Synchronous processing of tasks that users do not need to wait for
- No performance baseline — cannot detect degradation without a reference point
## Pillar 5: Cost Optimization
### Key Questions
- How do you implement cloud financial management?
- How do you govern usage and manage demand/supply?
- How do you evaluate new services for cost impact?
### Best Practices
- **Visibility**: Enable Cost Explorer, set up budgets and alerts, use cost allocation tags
- **Right-sizing**: Use Compute Optimizer recommendations, review utilization monthly
- **Pricing models**: Savings Plans for steady-state compute, Spot for fault-tolerant workloads, On-Demand for variable/unpredictable
- **Storage lifecycle**: S3 lifecycle policies to transition to Infrequent Access / Glacier
- **Eliminate waste**: Stop unused instances, delete unattached EBS volumes, remove unused Elastic IPs
### Common Anti-Patterns
- No cost allocation tagging (cannot attribute costs to teams or products)
- Running development environments 24/7 (schedule stop outside business hours)
- Paying On-Demand prices for predictable workloads (use Savings Plans)
- Unused Elastic IPs, idle load balancers, orphaned snapshots
## Pillar 6: Sustainability
### Key Questions
- How do you select regions to minimize carbon impact?
- How do you minimize resources required for your workload?
### Best Practices
- **Efficient resources**: Use Graviton (ARM) processors — better performance per watt
- **Scale to demand**: Auto-scale down during low traffic, schedule non-production shutdowns
- **Managed services**: Serverless and managed services optimize resource utilization across customers
- **Data management**: Delete unnecessary data, use appropriate storage tiers, compress data
## Well-Architected Review Process
### When to Conduct
- Before production launch (mandatory)
- Quarterly for critical workloads
- After significant architectural changes
- When performance or cost issues arise
### Steps
1. Select the workload scope (one application or service)
2. Walk through each pillar's questions with the team
3. Identify high-risk issues (HRIs) and improvement opportunities
4. Prioritize remediation by risk and effort
5. Create action items with owners and deadlines
6. Re-review after remediation to verify closure
@@ -0,0 +1,115 @@
# Regulatory Frameworks
Overview of major compliance frameworks, their requirements, and practical implementation guidance for cloud-native applications on AWS.
## PCI-DSS (Payment Card Industry Data Security Standard)
**Applies to**: Any system that stores, processes, or transmits cardholder data.
**Key Requirements** (organized by the 12 requirements):
1. **Network security**: Use security groups and NACLs to segment the cardholder data environment (CDE). No public internet access to CDE resources.
2. **Default credentials**: Change all vendor-supplied defaults. Automate with hardened AMIs and container images.
3. **Protect stored data**: Encrypt cardholder data at rest with KMS (AES-256). Implement data retention and disposal policies.
4. **Encrypt transmission**: TLS 1.2+ for all data in transit. Enforce HTTPS at ALB/API Gateway.
5. **Anti-malware**: Use Amazon Inspector for vulnerability scanning. GuardDuty for threat detection.
6. **Secure systems**: Patch management via SSM Patch Manager. IaC security scanning in CI.
7. **Access control**: Least-privilege IAM policies. No shared credentials. Role-based access.
8. **Authentication**: MFA for all administrative access. Cognito or IAM Identity Center for user management.
9. **Physical security**: Inherited from AWS for cloud infrastructure. Document shared responsibility model.
10. **Logging and monitoring**: CloudTrail for API activity. CloudWatch Logs for application logs. Retain for 1 year minimum.
11. **Testing**: Quarterly vulnerability scans. Annual penetration testing. IaC scanning in CI.
12. **Security policy**: Document and maintain an information security policy. Review annually.
**Scope reduction**: Use tokenization (via a PCI-compliant payment processor like Stripe) to minimize the CDE footprint. If you never touch raw card numbers, most PCI requirements do not apply.
## HIPAA (Health Insurance Portability and Accountability Act)
**Applies to**: Organizations handling Protected Health Information (PHI) in the US healthcare context.
**Key Requirements**:
- **BAA (Business Associate Agreement)**: Required with AWS before storing PHI. AWS offers BAAs for eligible services.
- **HIPAA-eligible services only**: Not all AWS services are HIPAA-eligible. Verify each service on the AWS HIPAA page.
- **Encryption**: PHI must be encrypted at rest (KMS) and in transit (TLS). This satisfies the Safe Harbor provision.
- **Access controls**: Role-based access to PHI. Audit all access with CloudTrail.
- **Audit trail**: Log all access to, creation of, and modification of PHI. Retain logs for 6 years.
- **Minimum necessary**: Only access and expose the minimum PHI required for the specific purpose.
- **Breach notification**: Notify affected individuals within 60 days of discovering a breach. Report to HHS.
**Architecture considerations**: Isolate PHI in dedicated accounts or VPCs. Use AWS PrivateLink to avoid PHI traversing the public internet. Tag all resources containing PHI for governance.
## SOC 2 (Service Organization Control 2)
**Applies to**: Service providers that store or process customer data. Increasingly expected by enterprise customers.
**Type I vs Type II**:
- **Type I**: Point-in-time assessment. "Controls are properly designed as of a specific date." Faster to achieve.
- **Type II**: Assessment over a period (usually 6-12 months). "Controls operated effectively during the review period." More rigorous and more valued.
**Trust Service Criteria**:
1. **Security** (required): Protection against unauthorized access. Covers firewalls, access controls, encryption, monitoring.
2. **Availability**: System is operational and accessible per SLA commitments.
3. **Processing Integrity**: System processing is complete, valid, accurate, and timely.
4. **Confidentiality**: Information designated as confidential is protected.
5. **Privacy**: Personal information is collected, used, retained, and disposed of per the privacy notice.
**Practical implementation**: Use AWS Config rules to continuously evaluate compliance. Automate evidence collection with Config conformance packs. Use Security Hub for centralized findings.
## GDPR (General Data Protection Regulation)
**Applies to**: Any organization processing personal data of EU/EEA residents, regardless of where the organization is located.
**Core Principles**:
- **Lawfulness**: Process data only with a legal basis (consent, contract, legitimate interest, legal obligation).
- **Purpose limitation**: Collect data only for specified, explicit purposes.
- **Data minimization**: Process only the data necessary for the stated purpose.
- **Accuracy**: Keep personal data accurate and up to date.
- **Storage limitation**: Retain data only as long as necessary. Define and enforce retention policies.
- **Integrity and confidentiality**: Protect data with appropriate security measures.
**Data Subject Rights**:
- Right of access (provide a copy of their data)
- Right to rectification (correct inaccurate data)
- Right to erasure ("right to be forgotten")
- Right to data portability (export in machine-readable format)
- Right to object to processing
**Technical implementation**: Build data subject access request (DSAR) automation. Implement soft-delete with configurable retention. Use DynamoDB TTL or S3 lifecycle policies for automatic data expiry. Tag PII data stores for governance.
## Data Residency and Sovereignty
- **Data residency**: Data must be stored within a specific geographic region. Choose AWS regions accordingly (eu-west-1 for EU, ap-southeast-2 for Australia).
- **Data sovereignty**: Data is subject to the laws of the country where it is stored.
- Use AWS Organizations SCPs to restrict resource creation to approved regions.
- Configure S3 bucket policies and DynamoDB table locations to enforce residency.
- Cross-region replication must respect residency requirements; do not replicate to unapproved regions.
- Document data flows across regions for compliance audits.
## Privacy Impact Assessment (PIA)
Conduct a PIA when introducing a new system or significantly changing data processing:
1. **Describe the processing**: What data, from whom, for what purpose, how long retained.
2. **Assess necessity**: Is the processing proportionate to the goal? Could the same outcome be achieved with less data?
3. **Identify risks**: Unauthorized access, accidental disclosure, data loss, function creep.
4. **Define mitigations**: Encryption, access controls, anonymization, pseudonymization, retention limits.
5. **Document and review**: Record the assessment. Review when processing changes or annually.
## Compliance-as-Code Patterns
Automate compliance verification rather than relying on manual audits:
- **AWS Config Rules**: Continuously evaluate resource configurations against compliance requirements (encrypted-volumes, restricted-ssh, mfa-enabled-for-iam).
- **Config Conformance Packs**: Pre-built rule sets for PCI-DSS, HIPAA, SOC2. Deploy via CloudFormation.
- **Security Hub**: Aggregate findings from Config, GuardDuty, Inspector, and third-party tools. Score against compliance frameworks.
- **cdk-nag**: Enforce compliance rules at synthesis time in CDK. Use AwsSolutionsChecks, NIST80053R5Checks, HIPAASecurityChecks, or PCIDSS321Checks.
- **Custom Config Rules**: Write Lambda-backed rules for organization-specific requirements.
## Audit Trail Requirements
Across all frameworks, a robust audit trail is a common requirement:
- **CloudTrail**: Enable in all regions, all accounts. Send to a centralized, immutable S3 bucket with Object Lock.
- **Application-level audit logs**: Log who did what, when, to which resource. Include the authentication context (user ID, role, IP).
- **Integrity protection**: Use CloudTrail log file integrity validation. S3 Object Lock for immutability.
- **Retention**: PCI-DSS: 1 year. HIPAA: 6 years. SOC2: per policy (typically 1-3 years). GDPR: as long as necessary for the purpose.
- **Access to logs**: Restrict log access to security and compliance roles. Log access to logs (meta-auditing) for sensitive environments.
@@ -0,0 +1,94 @@
# Composing a Workflow Plan
The composer's job is to fit the CEREMONY to the TASK: propose the minimum
viable workflow - the least sufficient EXECUTE set that still produces every
artifact the task's outcome depends on. Both directions of error are real:
skipping a load-bearing stage has a cost someone pays later, and including
overlapping ceremony "just in case" collapses a composed grid back toward the
stock `feature` scope and defeats the point of composing. Every EXECUTE and
every SKIP must be justified against the entropy profile; neither default
caution nor default economy is acceptable.
## How to read a task
- **Score before you select.** Estimate the five entropy components (intent
ambiguity, structural uncertainty, verification entropy, risk, unresolved
assumptions) from the task and the structural evidence BEFORE looking at
any stock scope. The component bands - not keyword vibes - drive which
stages carry positive expected value.
- **Incremental vs net-new.** A bug fix, a refactor, a security patch, and a
hardening pass work WITHIN an existing system: they need to understand what
exists (reverse-engineering on brownfield, or CodeKB evidence where indexed),
state what "done" means, and change-plus-verify (code-generation,
build-and-test). They do not need market-research, user-stories, or
domain-design - those discover and shape a product that already exists.
- **Net-new surface.** A new feature, product, or service needs the discovery
arc: intent-capture, scope-definition, then the inception design stages in
proportion to how much NEW structure it introduces.
- **Operational outcome.** Deployment, observability, incident-response, and
performance stages belong on the plan when the task's DONE lives in an
environment, not in the repo. A plan that builds but never ships closes no
operational task.
- **Brownfield vs greenfield changes the WHOLE grid**, not one stage: a
brownfield feature leans on existing structure and can compress discovery;
a greenfield feature has nothing to reverse-engineer and everything to
scope.
## Grid discipline
- Every required consume must have its producer on the EXECUTE set (the
validator enforces it; in-flight strict mode rejects). Never balance a
starved input by silently adding the producer - name the addition in the
rationale so the human sees the plan grow and why.
- Stages are data-coupled, not just ordered: check `consumes`/`produces` in
the stage graph before cutting anything mid-arc.
- Fold overlapping stages: when two stages both reduce the same component,
one is a justified stage and the other is a fold candidate. Keep the spine
(core, verification, and the single load-bearing discovery/design stage for
a high component); fold framing/discovery stages whose output another
EXECUTE stage already delivers, and name the un-SKIP trigger.
- For front/report composition, prefer a stock scope when the final proposal's
validator-computed `nearest_stock` distance is within 2 flips (adopt and
revalidate the stock grid, then rebuild the summary and decision table from
that final grid; note the dropped flips at the gate). The earlier mechanical
screen's distance is advisory and never overrides evidence-driven folds. A
custom scope is maintenance surface the user owns forever. A human edit to an
adopted stock grid converts it to custom so the edit has a persistence path.
When no stock scope fits the final proposal, synthesize - do not force a bad
match.
- In-flight recomposition never adopts a stock scope. Preserve the running
workflow's scope, depth, and frozen actions, then return only the strict-
validated pending delta as exact `changes.skip` / `changes.add` arrays for
the conductor's `recompose` command.
## Change Control
Every proposal names ONE Change Control value with a one-line rationale. The
value decides what happens when an input changes after the human approved or
confirmed something: `strict` reopens that approval; `relaxed` records the
change once, tells the human in one line, and continues. It never removes a
gate, so it is a question of how much the team wants to be asked again, not of
how much is checked.
- A matched stock scope carries its own default (`change_control:` in the
scope file; the shipped defaults are strict on enterprise, security-patch,
and infra, relaxed everywhere else). Adopt it and say so.
- For a custom grid, read the entropy profile the same way the grid was read:
high risk or verification entropy, regulated work, or several people sharing
the approvals point to strict; a spike, a fix, or a solo run where every
changed file would otherwise mean another approval points to relaxed.
- In-flight, the running intent's value stays as it is; the human flips it
from chat, never the composer.
- The human sees the value as its own gate row and can flip it before
approving. A memory layer that declares strict wins over any proposal; the
validator and the intent-create command both refuse a relaxed value under it.
## Rationale quality
The gate is only as good as the rationale. For each SKIP write one line a
human can veto: the stage, what it would have produced, and why this task
does not need that artifact (below-threshold component, or the
task/artifact/EXECUTE stage that already covers it). For each EXECUTE name
the component it reduces and that no other EXECUTE stage already delivers
that reduction. "Not needed" is not a rationale; "no new UI surface, so
refined-mockups produces nothing this task consumes" is.
@@ -0,0 +1,99 @@
# Mob Programming Guide
Practical guidance for effective mob programming (ensemble programming) as a team practice for knowledge sharing and high-quality delivery.
## What is Mob Programming?
Mob programming is the practice of the whole team working on the same thing, at the same time, in the same space (physical or virtual), at the same computer. It extends pair programming to the entire team.
Core principle: "All the brilliant minds working on the same thing, at the same time, in the same space, and at the same computer." — Woody Zuill
## Roles
### Driver
- The person at the keyboard. Types what the navigators direct.
- Does NOT make design decisions or solve problems independently.
- Focuses on translating spoken intent into code. Asks for clarification when directions are unclear.
- Think of the driver as a "smart input device" — capable and knowledgeable, but acting on group direction.
### Navigator(s)
- The rest of the team. They think, discuss, and direct the driver.
- One primary navigator speaks at a time to avoid overwhelming the driver.
- Navigators discuss approach, spot issues, suggest improvements, and think ahead.
- Different navigators bring different expertise: one thinks about design, another about edge cases, another about testing.
### Facilitator (optional, recommended for new mobs)
- Keeps the session on track. Manages rotation timer. Ensures everyone participates.
- Watches for dominant voices and draws quieter members into the conversation.
- Not a permanent role; rotate or remove once the team is comfortable with the practice.
## Rotation Cadence
- **Recommended interval**: 10-15 minutes per driver rotation.
- Use a timer (mobti.me, mob.sh CLI tool, or a simple kitchen timer).
- Everyone rotates through the driver role. No opt-outs; the practice only works with full participation.
- When the timer sounds, the current driver moves out, the next person in the rotation moves to the keyboard.
- The transition should be seamless: do not wait for a "good stopping point." Forcing handoff mid-thought builds shared understanding.
## Remote Mob Tooling
- **Screen sharing**: VS Code Live Share, JetBrains Code With Me, or plain screen share with remote control.
- **mob.sh**: CLI tool that automates git handoff. `mob start` creates a WIP branch; `mob next` commits and pushes for the next driver to pull. `mob done` squashes to a clean commit.
- **Timer tools**: mobti.me (web-based), mob.sh built-in timer, Cuckoo.team.
- **Communication**: Keep a persistent video call open. Audio quality matters more than video quality. Use a good microphone.
- **Shared notes**: Keep a shared document or whiteboard for parking lot items, decisions, and action items.
## When to Mob vs Pair vs Solo
| Situation | Recommended Practice |
|-----------|---------------------|
| New team member onboarding | Mob — fastest knowledge transfer |
| Complex design decision | Mob — multiple perspectives needed |
| Unfamiliar technology or domain | Mob — collective learning |
| Well-understood, repetitive work | Solo — mobbing adds overhead |
| Focused deep work (research, investigation) | Solo or pair — mob is too noisy |
| Code review backlog growing | Mob — eliminates the need for async review |
| Cross-team knowledge is siloed | Mob — breaks down silos |
| Time-sensitive bug fix | Pair or mob — faster diagnosis with multiple minds |
Mobbing is most valuable when uncertainty is high, knowledge needs to be shared, or quality matters more than raw throughput.
## Mob Session Facilitation
### Starting a Session
1. Agree on the goal: "By the end of this session, we want to have X."
2. Set the rotation timer.
3. Establish ground rules: respect the driver, one navigator speaks at a time, take breaks every 60-90 minutes.
4. Pull up all relevant context: tickets, design docs, existing code.
### During the Session
- If the mob gets stuck, take 5 minutes for silent individual research, then reconvene.
- Park tangential discussions on a visible "parking lot" list; address them later.
- If energy drops, take a break. A tired mob produces worse code than an individual.
- Celebrate small wins: passing tests, completing a feature, resolving a tricky bug.
### Ending a Session
- Commit and push all work (use `mob done` for a clean commit).
- Spend 5 minutes reviewing what was accomplished and what is left.
- Note any parking lot items that need follow-up.
## Knowledge Transfer Through Mobbing
- Mobbing is the fastest way to spread knowledge across a team. Every team member sees every decision in real time.
- New team members become productive faster because they absorb codebase knowledge, team conventions, and domain context simultaneously.
- Reduces bus factor to near zero: if one person leaves, the rest of the team has full context.
- Eliminates asynchronous code review: the review happens live, during development. Code is reviewed by the entire team before it is committed.
## Mob Retrospectives
After running mob sessions for 1-2 weeks, hold a retrospective specifically about the practice:
- **What is working well?** (knowledge sharing, fewer bugs, faster onboarding)
- **What is frustrating?** (rotation too fast/slow, some people dominating, fatigue)
- **What should we experiment with?** (different rotation time, mob only for complex work, include stakeholders)
Common adjustments:
- Increase rotation time if transitions feel disruptive (try 15-20 minutes).
- Decrease rotation time if the driver disengages or dominates (try 7-10 minutes).
- Mob for half the day and solo for the other half if energy is a concern.
- Use strong-style pairing rule: "For an idea to go from your head into the computer, it must go through someone else's hands."
@@ -0,0 +1,80 @@
# Team Topologies
Organising teams for fast flow of change using the Team Topologies framework by Matthew Skelton and Manuel Pais.
## The Four Fundamental Team Types
### 1. Stream-Aligned Team
- **Purpose**: Delivers value along a single stream of work (a product, a feature set, a user journey, or a business domain).
- **Characteristics**: Cross-functional (dev, test, ops, UX). Owns the full lifecycle from ideation to production. Has clear ownership boundaries.
- **Size**: 5-9 people (two-pizza rule). Enough to own a meaningful slice of the product without excessive coordination.
- **This is the primary team type.** Most teams in the organisation should be stream-aligned. Other team types exist to reduce the cognitive load on stream-aligned teams.
### 2. Platform Team
- **Purpose**: Provides internal services that accelerate stream-aligned teams. Reduces cognitive load by abstracting away infrastructure complexity.
- **Examples**: Internal developer platform (IDP), CI/CD pipeline team, observability platform, shared authentication service.
- **Operates as a product team**: Treats stream-aligned teams as customers. Publishes a clear API/interface. Prioritises usability and self-service.
- **Anti-pattern**: A platform team that requires tickets and manual intervention is a bottleneck, not a platform.
### 3. Enabling Team
- **Purpose**: Helps stream-aligned teams acquire new capabilities. Coaches, mentors, and researches — does not build features.
- **Examples**: Cloud adoption team helping migrate from on-premise. Security enablement team coaching secure coding practices. SRE team teaching observability patterns.
- **Time-boxed engagement**: Works with a stream-aligned team for weeks or months, then moves on. Success means the stream-aligned team no longer needs the enabling team.
### 4. Complicated-Subsystem Team
- **Purpose**: Owns a component that requires deep specialist knowledge that most stream-aligned teams cannot reasonably maintain.
- **Examples**: ML model training pipeline, video codec optimization, cryptography module, real-time data processing engine.
- **Rare**: Only create this team type when the subsystem's complexity truly justifies specialist ownership. Over-use leads to silos.
## Three Interaction Modes
| Mode | Description | When to Use |
|------|-------------|-------------|
| **Collaboration** | Two teams work closely together on a shared goal. High communication bandwidth. | Discovery phases, building a new capability, exploring uncertainty. Time-box to avoid permanent coupling. |
| **X-as-a-Service** | One team provides a service; the other consumes it via a well-defined API. Low communication overhead. | Stable, well-understood capabilities. The providing team's interface is mature. |
| **Facilitating** | One team (typically enabling) helps another team learn or adopt a new practice. | Skill transfer, technology adoption, practice improvement. |
Interaction modes should evolve over time. A platform team might start in collaboration mode with a stream-aligned team and transition to X-as-a-service once the interface stabilises.
## Cognitive Load Assessment
Cognitive load is the primary constraint on team effectiveness. Three types:
- **Intrinsic**: Complexity of the domain itself (financial regulations, distributed systems).
- **Extraneous**: Unnecessary complexity from tooling, process, or poor documentation. Reduce this.
- **Germane**: Productive learning related to the domain. Increase this.
**Assessment questions for each team**:
1. How many services/components does this team own? (If > 3-5 significant services, the team is overloaded.)
2. How many different technology stacks must the team maintain?
3. How much time is spent on operational toil vs feature development?
4. How often does the team need to coordinate with other teams to deliver?
5. Can a new team member become productive within 2-4 weeks?
If cognitive load is too high, split the team's responsibilities, create a platform team to absorb shared concerns, or simplify the architecture.
## Team API Concept
Each team should publish a "Team API" that describes:
- **What the team owns**: Services, data stores, APIs, domains.
- **How to interact**: Preferred communication channels, office hours, request processes.
- **What the team provides**: Interfaces, SLOs, documentation, support expectations.
- **What the team needs**: Dependencies on other teams, expected SLOs from dependencies.
The Team API makes boundaries explicit and reduces ad-hoc interruptions.
## Conway's Law Implications
"Organizations which design systems are constrained to produce designs which are copies of the communication structures of these organizations." — Melvin Conway
**Practical implications**:
- If you want a microservices architecture, organise teams around services. A monolithic team structure will produce a monolith.
- If two teams must collaborate to deploy a feature, the architecture has an implicit coupling that should be addressed.
- Use the "Inverse Conway Manoeuvre": design the team structure to match the desired architecture, and the architecture will follow.
## Team Sizing — The Two-Pizza Rule
- A team should be small enough that two pizzas can feed it (5-9 people).
- Below 5: insufficient breadth of skills; high bus factor risk.
- Above 9: communication overhead grows quadratically; decision-making slows.
- If a team is growing beyond 9, look for a natural boundary to split along (a subdomain, a component, a user journey).
@@ -0,0 +1,147 @@
# Workflow Planning Guide
Domain-specific guidance for the Workflow Planning stage. Use this alongside `product-guide.md` when leading execution plan creation.
## Stage Configuration Heuristics
Derive stage configuration from the work breakdown analysis. For each conditional stage, evaluate whether it adds value based on the identified work streams, their complexity, and dependencies.
### INCEPTION stages
| Stage | EXECUTE when | SKIP when |
|-------|-------------|-----------|
| Domain Design | Work streams introduce new components/services, new architectural boundaries, greenfield projects | All streams modify existing components only, no new service boundaries |
| Units Generation | Multiple independent work streams, cross-cutting concerns requiring sequenced delivery | Single stream or tightly coupled streams that form one natural unit |
| Contract Design | More than one unit must integrate, or a unit exposes a public/external API to formalise before parallel build | Single self-contained unit with no inter-unit boundaries and no external API |
### CONSTRUCTION stages (per-unit)
| Stage | EXECUTE when | SKIP when |
|-------|-------------|-----------|
| Functional Design | Streams involve complex business logic, state machines, multi-step workflows, domain modeling | Simple CRUD, config changes, straightforward data transformations |
| NFR Requirements | Streams handle security-sensitive data, performance SLAs, public-facing APIs, regulatory compliance | Internal tools, prototypes, low-risk utility functions |
| NFR Design | NFR Requirements produced non-trivial requirements | NFR Requirements skipped or produced only basic constraints |
| Infrastructure Design | Streams require new deployment targets, CI/CD changes, infrastructure-as-code | Existing infrastructure unchanged, deploying to established pipeline |
### Configuration rationale format
For each stage decision, tie the rationale to specific work streams:
- "EXECUTE — Streams 1 and 3 introduce new service boundaries requiring architectural design"
- "SKIP — All streams modify existing components within established architecture"
## Work Stream Identification Patterns
### Grouping strategies
- **By domain area**: Group requirements/stories that share domain entities and business rules
- **By user persona**: Group requirements serving the same user type
- **By dependency chain**: Group requirements where one enables another
- **By risk profile**: Isolate high-risk work into its own stream for focused attention
- **By delivery boundary**: Group work that can be independently delivered and tested
### Stream sizing guidance
- **Simple projects** (1-2 streams): Single feature additions, bug fixes, focused refactoring
- **Standard projects** (2-4 streams): Multi-feature work with some cross-cutting concerns
- **Complex projects** (4-6 streams): Distributed changes, multiple integration points, significant architectural work
### Sequencing strategies
1. **Foundation first**: Infrastructure and shared services before dependent features
2. **High-risk early**: Tackle uncertainty before investing in dependent work
3. **Value delivery**: Arrange so partial delivery still provides user value
4. **Test isolation**: Each stream should be independently testable where possible
5. **Critical path optimization**: Identify the longest dependency chain and prioritize unblocking it
## Economic vs topological sequencing (for Bolt Planning in Stage 2.9)
Unit dependency analysis (Stage 2.7) produces the DAG — topological order falls out of it mechanically. That's geometry: what the system is.
Bolt sequencing (Stage 2.9) is different work. It chooses a path through the DAG weighted by human value judgment — which Bolt ships first, which proves what, which surfaces the biggest risk early. AI can topologically sort; it cannot decide what validates the market hypothesis fastest.
Per the canonical Glossary (`stage-protocol.md` Terminology), a **Bolt** is the planned Construction delivery slice from 2.9: one or more Units with a Definition of Done, a confidence hypothesis, and ownership. Bolts are not MMFs and not sprints.
Heuristics for Bolt sequencing:
- **Walking skeleton first** (Cockburn, *Crystal Clear*) — the first Bolt is a minimal end-to-end implementation that proves the architecture works, before adding features.
- **WSJF / Cost of Delay ÷ Duration** (Reinertsen, *Principles of Product Development Flow*; SAFe) — order Bolts by (value + time criticality + risk reduction) divided by job size.
- **Risk-first** (Boehm, Spiral Model) — sequence the highest-uncertainty Bolts early so decisions are calibrated before dependent work commits.
- **Value-first** — ship Bolts in value order when risk is low and value delivery is the dominant constraint.
The chosen heuristic is captured in `risk-and-sequencing-rationale.md`, alongside any deviation from 2.7's topological order.
## Execution Plan Structure
Every execution plan MUST include these sections:
1. **Work Streams** — identified streams with scope, deliverables, complexity, dependencies, requirements coverage, review expertise
2. **Implementation Sequence** — ordered stream execution with critical path
3. **Detailed Analysis Summary** — scope metrics, change impact, component relationships
4. **Risk Assessment** — risk register with likelihood, impact, and mitigation
5. **Workflow Visualization** — Mermaid flowchart of stage execution flow
6. **Stage Configuration** — checkbox list of stages with EXECUTE/SKIP decisions and rationale tied to work streams
7. **Success Criteria** — measurable outcomes for project completion
Optional sections (include when applicable):
- **Transformation Scope** — for brownfield projects with significant refactoring
- **Package Change Sequence** — for multi-unit projects with dependency ordering
- **Multi-Module Coordination** — for brownfield projects touching multiple packages
## Risk Assessment Criteria
### Severity Levels
| Level | Description | Indicators | Example |
|-------|-------------|------------|---------|
| **Low** | Well-understood, minimal dependencies | Standard patterns, established tech, isolated changes | Adding a new REST endpoint to an existing API |
| **Medium** | Some unknowns, moderate dependencies | New library adoption, moderate cross-component impact | Integrating a third-party auth provider |
| **High** | Significant unknowns, complex dependencies | New technology, data migration, multiple integration points | Migrating from SQL to NoSQL for a core domain |
| **Critical** | Architectural changes, breaking changes | Fundamental pattern changes, data schema overhaul, API contract changes | Rewriting monolith services into microservices |
### Risk Documentation Pattern
For each identified risk, document:
- **Risk**: What could go wrong
- **Likelihood**: Low / Medium / High
- **Impact**: Low / Medium / High / Critical
- **Mitigation**: Specific actions to reduce likelihood or impact
## Unit Decomposition Heuristics
### When to use single-unit delivery
- Fewer than 5 user stories
- All stories share the same components
- No independent deploy/test boundaries
- Simple feature addition or bug fix
### When to use multi-unit delivery
- 5+ user stories spanning different domains
- Independent feature groups that can be delivered and tested separately
- Different risk profiles across feature groups (ship low-risk first)
- Cross-cutting concerns (e.g., auth, logging) that should be built before dependent features
### Unit ordering principles
1. **Foundation first**: Infrastructure and shared services before dependent features
2. **High-risk early**: Tackle uncertainty before investing in dependent work
3. **Value delivery**: Arrange so partial delivery still provides user value
4. **Test isolation**: Each unit should be independently testable
## Depth Calibration
### Simple project indicators
- Single page or single API endpoint
- No external integrations
- Single user role
- Straightforward CRUD operations
- Internal tool or prototype
### Standard project indicators
- Multi-page application or multi-endpoint API
- 1-3 external integrations
- 2-4 user roles with different permissions
- Some business logic beyond CRUD
- Production-grade with moderate traffic expectations
### Complex project indicators
- Distributed system or microservice architecture
- 4+ external integrations or real-time data flows
- Complex authorization model (RBAC, ABAC, multi-tenancy)
- Domain-specific algorithms, state machines, or workflow engines
- High availability requirements, data migration, regulatory compliance
@@ -0,0 +1,120 @@
# Accessibility: WCAG 2.1 AA Guide
## Purpose
Ensure digital products are usable by people with diverse abilities. WCAG 2.1 Level AA is the standard target for most applications and is legally required in many jurisdictions.
## Four Principles (POUR)
### 1. Perceivable
Information and UI components must be presentable in ways users can perceive.
**Text Alternatives**
- All non-decorative images must have descriptive `alt` text
- Complex images (charts, diagrams) need long descriptions
- Decorative images use `alt=""` (empty) to be ignored by screen readers
**Color Contrast**
- Normal text (< 18px): minimum 4.5:1 contrast ratio against background
- Large text (>= 18px bold or >= 24px regular): minimum 3:1 contrast ratio
- UI components and graphical objects: minimum 3:1 contrast ratio
- Never use color alone to convey meaning (add icons, text labels, or patterns)
**Media**
- Video must have captions (synchronized with audio)
- Audio-only content needs text transcripts
- No content that flashes more than 3 times per second
### 2. Operable
UI components and navigation must be operable by all users.
**Keyboard Navigation**
- All functionality must be accessible via keyboard alone
- Visible focus indicator on every interactive element (minimum 2px outline, 3:1 contrast)
- Logical tab order following visual layout (left-to-right, top-to-bottom)
- No keyboard traps — users must be able to navigate away from any component
- Skip-to-content link as the first focusable element on each page
**Keyboard Patterns by Component**
| Component | Keys |
|-----------|------|
| Buttons | Enter or Space to activate |
| Links | Enter to follow |
| Checkboxes | Space to toggle |
| Radio buttons | Arrow keys to move between options |
| Tabs | Arrow keys to switch, Tab to enter/exit tab panel |
| Modals | Escape to close, trap focus within modal |
| Dropdowns | Arrow keys to navigate, Enter to select, Escape to close |
**Timing**
- No time limits on interactions, or provide option to extend/disable
- Auto-updating content can be paused, stopped, or hidden
### 3. Understandable
Information and UI operation must be understandable.
**Readability**
- Page language declared in HTML (`lang` attribute)
- Consistent navigation across pages
- Consistent identification of repeated components
**Predictability**
- No unexpected context changes on focus or input
- Form submissions require explicit user action (button click, not auto-submit)
- Navigation order is consistent across pages
**Error Handling**
- Input errors are identified and described in text (not just red borders)
- Labels and instructions are provided before input fields
- Error suggestions offer specific correction guidance
- Important submissions (financial, legal) are reversible, verified, or confirmed
### 4. Robust
Content must be robust enough for diverse user agents and assistive technologies.
**Markup**
- Valid, well-structured HTML with proper nesting
- Unique IDs throughout the page
- Complete start and end tags
## ARIA (Accessible Rich Internet Applications)
### When to Use ARIA
- **First rule**: Use native HTML elements whenever possible. A `<button>` is always better than `<div role="button">`
- Use ARIA only when native HTML cannot express the semantics
### Essential ARIA Attributes
- `role` — Defines what the element is (e.g., `dialog`, `alert`, `tabpanel`, `navigation`)
- `aria-label` — Provides accessible name when visible text is insufficient
- `aria-labelledby` — Points to another element that labels this one
- `aria-describedby` — Points to element providing additional description
- `aria-expanded` — Indicates if a collapsible section is open (true/false)
- `aria-hidden="true"` — Hides decorative elements from screen readers
- `aria-live="polite"` — Announces dynamic content changes (toast messages, status updates)
- `aria-required="true"` — Marks required form fields
### ARIA Landmarks
- `role="banner"` or `<header>` — Site-wide header
- `role="navigation"` or `<nav>` — Navigation blocks
- `role="main"` or `<main>` — Primary content area
- `role="complementary"` or `<aside>` — Supporting content
- `role="contentinfo"` or `<footer>` — Site-wide footer
## Common Failures and Fixes
| Failure | Impact | Fix |
|---------|--------|-----|
| Missing alt text on images | Screen readers say "image" with no context | Add descriptive alt text or `alt=""` for decorative |
| Low color contrast | Unreadable for low vision users | Use contrast checker, meet 4.5:1 minimum |
| No focus indicators | Keyboard users cannot see where they are | Add visible `:focus` styles, never use `outline: none` |
| Form fields without labels | Screen readers cannot identify inputs | Associate `<label>` with every `<input>` via `for`/`id` |
| Auto-playing media | Disorienting, interferes with screen readers | Require user action to play, provide pause/stop |
| Mouse-only interactions | Keyboard/switch users cannot operate | Add keyboard event handlers for all mouse interactions |
| Missing heading hierarchy | Navigation by headings fails | Use h1-h6 in logical order, never skip levels |
| Dynamic content without announcements | Screen readers miss updates | Use `aria-live` regions for status messages |
## Testing Approach
1. **Automated scan**: axe DevTools, Lighthouse accessibility audit (catches ~30% of issues)
2. **Keyboard testing**: Unplug the mouse and navigate the entire application
3. **Screen reader testing**: Test with VoiceOver (macOS), NVDA (Windows), or TalkBack (Android)
4. **Zoom testing**: Verify layout at 200% and 400% browser zoom
5. **Color testing**: Verify with simulated color blindness (protanopia, deuteranopia, tritanopia)
@@ -0,0 +1,61 @@
# Component Specification Template
Use this template for component-level specifications in `interaction-spec.md` (Stage 1.5 Refined Mockups) and any stage requiring detailed UI component definitions.
---
## [Component Name]
| Field | Value |
|---|---|
| Component | [name] |
| Description | [one-line purpose] |
| Category | [input / display / layout / navigation / feedback] |
### States
| State | Description | Trigger |
|---|---|---|
| default | Initial render state | page load |
| hover | Cursor over element | mouseover |
| focus | Keyboard focus | Tab key / click |
| disabled | Non-interactive | prop disabled=true |
| loading | Async operation pending | async op in progress |
| error | Validation or system error | validation failure |
| empty | No data to display | no data |
### Props / Inputs
| Prop | Type | Required | Default | Description |
|---|---|---|---|---|
| [prop-name] | [string \| boolean \| number \| object] | [yes/no] | [value or —] | [what it controls] |
### Responsive Behaviour
| Breakpoint | Behaviour |
|---|---|
| mobile (<768px) | [layout/visibility changes] |
| tablet (768–1024px) | [layout/visibility changes] |
| desktop (>1024px) | [default layout] |
### Accessibility
| Requirement | Implementation |
|---|---|
| ARIA role | [role — e.g. button, listbox, dialog] |
| Keyboard interaction | [Tab to focus, Enter/Space to activate, Escape to dismiss] |
| Label / aria-label | [visible label, aria-label, or aria-labelledby approach] |
| Contrast ratio | WCAG AA (4.5:1 text, 3:1 UI components) |
| Screen reader | [what is announced and when] |
| Focus management | [where focus goes on open/close/activate] |
### Usage Example
```
<ComponentName
prop="value"
onAction={handler}
/>
```
---
@@ -0,0 +1,146 @@
# Interaction Design Patterns
## Purpose
Reusable solutions to common UI interaction problems. Applying established patterns reduces user learning curve and development effort.
## Navigation Patterns
### Top Navigation Bar
- Best for: Applications with 3-7 top-level sections
- Include: Logo (home link), primary nav items, user menu, search
- On mobile: Collapse to hamburger menu or bottom tab bar
### Side Navigation
- Best for: Applications with many sections, deep hierarchies, or admin interfaces
- Collapsible to icons-only for more content space
- Active section should be visually highlighted
- Support nested items with expand/collapse
### Breadcrumbs
- Best for: Deep hierarchies (e-commerce, file systems, documentation)
- Show the path from root to current page
- Each segment is a clickable link except the current page
- Do not use breadcrumbs as the only navigation method
### Bottom Tab Bar (Mobile)
- Best for: Mobile apps with 3-5 primary sections
- Maximum 5 tabs; more than 5 requires a "More" overflow
- Active tab uses filled icon and label; inactive tabs use outlined icons
## Form Patterns
### Inline Validation
- Validate on blur (when the user leaves the field), not on every keystroke
- Show success state for valid fields to build confidence
- Place error messages directly below the relevant field
- Use specific error messages: "Password must be at least 8 characters" not "Invalid input"
### Multi-Step Forms (Wizards)
- Show progress indicator (step 1 of 4) with step labels
- Allow backward navigation to review previous steps
- Save progress between steps (do not lose data on back-navigation)
- Final step shows a summary for review before submission
- Keep each step focused on one logical group of inputs
### Autosave
- Save drafts automatically at intervals or on field change
- Show save status clearly: "Saved", "Saving...", "Unsaved changes"
- Provide explicit save/discard actions for critical data
## Modal and Dialog Patterns
### When to Use Modals
- Confirming destructive actions ("Delete this item?")
- Collecting small amounts of focused input (rename, quick settings)
- Displaying critical alerts that require acknowledgment
### When NOT to Use Modals
- Displaying large amounts of content (use a new page instead)
- Nested modals (modal opening another modal — always avoid)
- Optional information (use inline expansion or tooltips)
### Modal Implementation Rules
- Trap keyboard focus inside the modal while open
- Close on Escape key press
- Close on overlay/backdrop click (except for critical confirmations)
- Return focus to the trigger element when closed
- Prevent background scrolling while modal is open
## Progressive Disclosure
### Pattern
Show only essential information initially; reveal detail on demand.
### Applications
- **Accordion sections**: Collapse secondary content; expand on click
- **"Show more" links**: Truncate long lists/text with option to expand
- **Advanced settings**: Hide behind a "Show advanced options" toggle
- **Contextual help**: Show tips/explanations via info icons or tooltips, not inline clutter
### Rule
Every screen should have a clear primary action. If users are overwhelmed, you are showing too much at once.
## Infinite Scroll vs Pagination
### Infinite Scroll
- Best for: Social feeds, media galleries, content discovery
- Show loading indicator at bottom when fetching more
- Provide "Back to top" button after scrolling
- Caution: Breaks browser back button, makes footer unreachable, loses scroll position
### Pagination
- Best for: Search results, data tables, e-commerce listings
- Show total count and current position ("Showing 1-20 of 347")
- Include: Previous, Next, first/last page, and 2-3 surrounding page numbers
- Preserve filter/sort state across page changes
## Drag and Drop
### When Appropriate
- Reordering lists, kanban boards, file uploads, layout builders
- Always provide a non-drag alternative (move up/down buttons, keyboard shortcuts)
### Implementation
- Show a grab cursor on hover of draggable items
- Provide a clear visual drop target (highlighted zone, insertion line)
- Show a ghost/preview of the dragged item
- Support undo immediately after drop (Ctrl+Z or undo toast)
## Micro-Interactions
### Definition
Small, single-purpose animations or feedback moments that make the interface feel responsive.
### Key Micro-Interactions
- **Button feedback**: Subtle press/depress animation on click
- **Toggle transitions**: Smooth state change (on/off) with color shift
- **Success confirmation**: Brief checkmark animation after form submission
- **Skeleton loading**: Content-shaped placeholders that pulse while loading
- **Pull to refresh**: Resistance and spinner animation (mobile)
### Rules
- Keep animations under 300ms — longer feels sluggish
- Use easing (ease-out for entrances, ease-in for exits) — linear motion feels robotic
- Respect `prefers-reduced-motion` media query — disable animations for users who request it
## Error Prevention Patterns
- **Confirmation dialogs** for destructive actions (delete, overwrite, send)
- **Undo** instead of confirmation when possible (Gmail's "Undo send" is superior to "Are you sure?")
- **Constraints**: Disable invalid options rather than showing errors after selection
- **Defaults**: Pre-fill with sensible defaults to reduce input errors
- **Format hints**: Show expected format inline ("DD/MM/YYYY") not just in error messages
## Responsive Breakpoint Strategy
### Standard Breakpoints
- **Mobile**: 320px - 767px (single column, stacked layout)
- **Tablet**: 768px - 1023px (two columns, collapsible side nav)
- **Desktop**: 1024px - 1439px (full layout, side nav expanded)
- **Large desktop**: 1440px+ (max-width container, avoid stretching content beyond ~1200px)
### Design Approach
- Design mobile-first: start with the smallest screen, add complexity as space allows
- Use fluid grids and relative units (%, rem) not fixed pixels
- Test at breakpoint boundaries AND mid-points (avoid layout breaking at 900px between 768 and 1024)
- Touch targets: minimum 44x44px on mobile (Apple HIG), 48x48px (Material Design)
@@ -0,0 +1,139 @@
# UX Guide
## Nielsen's 10 Usability Heuristics (Applied)
Use these as a review checklist for every user-facing specification:
1. **Visibility of system status**: Show loading indicators, progress bars, success confirmations. Users must always know what is happening.
2. **Match between system and real world**: Use domain language the user understands. Avoid technical jargon in UI labels.
3. **User control and freedom**: Provide undo, cancel, and back. Never trap users in a flow without an exit.
4. **Consistency and standards**: Same action = same label = same position across all screens. Follow platform conventions.
5. **Error prevention**: Disable invalid actions, use type-appropriate inputs (date pickers, dropdowns), confirm destructive operations.
6. **Recognition rather than recall**: Show options, recent items, defaults. Minimize what users must remember between screens.
7. **Flexibility and efficiency of use**: Support keyboard shortcuts, bulk actions, and saved preferences for expert users without cluttering the novice experience.
8. **Aesthetic and minimalist design**: Every element must earn its place. Remove decorative elements that do not aid task completion.
9. **Help users recognize, diagnose, and recover from errors**: Error messages must say what went wrong, why, and what to do next. Never show raw error codes.
10. **Help and documentation**: Provide contextual help (tooltips, inline hints) at the point of need, not in a separate help section.
## WCAG 2.1 AA Key Requirements
### Perceivable
- Text contrast ratio: minimum 4.5:1 for normal text, 3:1 for large text (18px+ bold or 24px+)
- Non-text content has text alternatives (alt text, ARIA labels)
- Content does not rely solely on color to convey meaning (use icons, patterns, or text too)
- Media has captions or transcripts
### Operable
- All functionality available via keyboard (Tab, Enter, Space, Arrow keys, Escape)
- Focus order follows a logical reading sequence
- Focus indicators are visible (never `outline: none` without replacement)
- No content flashes more than 3 times per second
- Touch targets are minimum 44x44 CSS pixels
### Understandable
- Language of page is declared in HTML
- Form inputs have visible labels (not just placeholders)
- Error identification is specific ("Email is invalid" not "Error in field 3")
- Consistent navigation across pages
### Robust
- Valid HTML semantics (headings in order, lists for lists, tables for tabular data)
- ARIA roles used correctly (not overused; native HTML elements preferred)
- Content works across browsers and assistive technologies
## Interaction Patterns Reference
### Forms
- Labels above inputs (not beside, not inside as placeholder-only)
- Inline validation on blur, not on every keystroke
- Submit button disabled until required fields are valid (with visual explanation)
- Group related fields with fieldset/legend
- Mark optional fields, not required ones (most fields should be required)
### Data Tables
- Sortable columns with sort indicator (arrow direction)
- Filterable with clear filter indicators and reset option
- Pagination with page size selector and total count
- Row selection with bulk action toolbar
- Empty state message with guidance ("No results. Try adjusting your filters.")
### Navigation
- Primary navigation: persistent, max 7 items, current page highlighted
- Breadcrumbs for hierarchical content (3+ levels deep)
- Search: globally accessible, auto-suggest after 3 characters, recent searches shown
### Feedback & States
- **Loading**: Skeleton screens for initial load, spinners for actions (with timeout message after 5s)
- **Success**: Inline confirmation near the action, auto-dismiss after 5s, do not redirect immediately
- **Error**: Inline near the cause, red but with icon (not color-only), actionable message
- **Empty**: Illustration + explanation + primary action ("No items yet. Create your first item.")
- **Confirmation**: Required for delete, bulk operations, and irreversible actions. Include what will happen and an undo option if possible.
## User Flow Documentation Format
For each user flow, specify:
```
Flow: [name]
Persona: [who performs this]
Trigger: [what initiates the flow]
Steps:
1. [Screen/state] -> [user action] -> [system response]
2. [Screen/state] -> [user action] -> [system response]
...
Success outcome: [what the user sees when done]
Error paths:
- [condition] -> [error screen/message] -> [recovery action]
```
## Information Architecture
### Navigation Hierarchy Principles
- Maximum 3 levels of nesting for primary navigation
- Flat is better than deep — prefer broad categories with fewer sub-levels
- Every page must be reachable within 3 clicks from the home/dashboard
- Use progressive disclosure: show summary first, detail on demand
### Content Grouping Strategies
- Group by user task (what they want to do), not by system structure (how it is built)
- Card sorting reference: use open card sort for new IA, closed card sort to validate existing
- Related actions should be visually proximate (Gestalt principle of proximity)
### Labeling Taxonomy
- Labels must use the user's language, not internal jargon
- Consistent verb forms across navigation (all nouns or all verbs, not mixed)
- Test labels with 5+ representative users before finalizing
### Sitemap Structure Patterns
- Hub-and-spoke: central dashboard with links to feature areas (suits task-based apps)
- Hierarchical: nested tree (suits content-heavy sites, documentation)
- Sequential: linear flow (suits onboarding, checkout, wizards)
- Choose the pattern that matches the primary user workflow
## Responsive and Adaptive Design
### Breakpoint Strategy (Mobile-First)
- Design for smallest screen first, then enhance for larger screens
- Common breakpoints: 320px (mobile), 768px (tablet), 1024px (desktop), 1440px (large desktop)
- Content dictates breakpoints, not device names — add breakpoints where the layout breaks
### Layout Adaptation Patterns
- **Fluid**: Percentage-based widths, content reflows naturally (default approach)
- **Adaptive**: Distinct fixed layouts per breakpoint (use when fluid is insufficient)
- **Responsive**: Combination of fluid grids, flexible images, and media queries (recommended)
### Touch Target Sizing
- Minimum 44x44 CSS pixels for all interactive elements (WCAG 2.1 AA)
- 8px minimum spacing between adjacent touch targets
- Increase to 48x48 px for primary actions on mobile
### Content Priority Shifting
- Stack columns vertically on mobile (most important content first)
- Hide secondary navigation behind a menu icon on small screens
- Collapse data tables into card views on mobile
- Defer non-critical images and media on slow connections
### Performance Considerations for Mobile
- Target < 3s load time on 3G connections
- Lazy-load images and below-the-fold content
- Minimize JavaScript payload (< 200KB compressed for initial load)
- Use responsive images (srcset) to serve appropriately sized assets
@@ -0,0 +1,109 @@
# Wireframing Guide
## Purpose
Wireframes are visual blueprints that define layout, hierarchy, and interaction flow before visual design or development begins. They reduce rework by validating structure early and cheaply.
## Fidelity Progression
### Low-Fidelity (Sketches)
- **Tools**: Paper, whiteboard, basic drawing tools
- **When**: Initial ideation, stakeholder alignment, exploring multiple layouts quickly
- **Content**: Boxes and lines, placeholder text ("Lorem ipsum"), no color, no real data
- **Time per screen**: 5-15 minutes
- **Rule**: If you spend more than 15 minutes on a sketch, you are over-investing
### Mid-Fidelity (Structural Wireframes)
- **Tools**: Figma, Balsamiq, Excalidraw
- **When**: Defining content hierarchy, information architecture, navigation flow
- **Content**: Real labels, approximate spacing, grayscale, actual content structure
- **Time per screen**: 30-60 minutes
### High-Fidelity (Interactive Prototypes)
- **Tools**: Figma (prototyping mode), Framer
- **When**: User testing, developer handoff, complex interaction validation
- **Content**: Real copy, accurate spacing, clickable interactions, state transitions
- **Time per screen**: 2-4 hours
### Progression Rule
Start at the lowest fidelity that answers your current question. Do not jump to high-fidelity until low-fidelity concepts are validated.
## Layout Patterns
### F-Pattern (Content-Heavy Pages)
Users scan horizontally across the top, then down the left side, then across again. Use for:
- Article pages, search results, dashboards
- Place critical content in the top-left and along the left edge
### Z-Pattern (Marketing / Landing Pages)
Eye moves: top-left to top-right, diagonally to bottom-left, then to bottom-right. Use for:
- Landing pages, sign-up flows, simple layouts
- Place logo top-left, CTA top-right, key message bottom-right
### Card Layout
Grid of self-contained content units. Use for:
- Product catalogs, dashboards, media galleries
- Each card is independently scannable and actionable
### Split Screen
Two equal or weighted panels side by side. Use for:
- Comparison views, master-detail, editor-preview
## Component Library Basics
### Essential Components to Define Early
- **Navigation**: Top bar, side nav, breadcrumbs, tabs
- **Data display**: Tables, cards, lists, detail panels
- **Input**: Text fields, selects, checkboxes, date pickers, file upload
- **Feedback**: Alerts, toasts, progress bars, empty states
- **Actions**: Buttons (primary, secondary, destructive), links, menus
### Consistency Rules
- One primary action per screen section (single prominent button)
- Consistent placement of navigation and actions across all screens
- Uniform spacing scale (4px, 8px, 16px, 24px, 32px, 48px)
## Screen State Design
Every screen has multiple states. Wireframe ALL of them, not just the happy path.
### The Five States
1. **Empty State**
- First-time user with no data
- Include: illustration/icon, explanation of what will appear, clear CTA to create first item
- Never show a blank table or empty list with no guidance
2. **Loading State**
- Data is being fetched or processed
- Use skeleton screens (preferred) or spinners
- Show loading in context (inline), not as a full-page block
3. **Success / Populated State**
- Normal operation with real data
- This is the state most wireframes show — but it is only one of five
4. **Error State**
- Something went wrong (network error, validation failure, permission denied)
- Explain what happened in plain language, suggest recovery action
- Never show raw error codes or stack traces to users
5. **Partial / Edge State**
- Incomplete data, very long content, single item vs many items, max limits reached
- Test with: 0 items, 1 item, 5 items, 100 items, 10,000 items
- Test with: very short text, very long text, special characters, missing optional fields
## Wireframe Review Checklist
- [ ] All five screen states represented (empty, loading, success, error, partial)
- [ ] Navigation is consistent across all screens
- [ ] Content hierarchy is clear — most important information is most prominent
- [ ] Interactive elements are obviously clickable/tappable
- [ ] Mobile and desktop layouts considered (even if only one is wireframed in detail)
- [ ] Accessibility annotations present (tab order, heading levels, alt text notes)
- [ ] Edge cases documented (long names, missing data, permission variations)
## Common Mistakes
- Wireframing only the happy path with perfect data
- Using placeholder text that hides layout problems (real content is longer/shorter)
- Skipping mobile layout until development
- Treating wireframes as final design — they should invite feedback and iteration
- Not annotating interaction behavior (what happens on click, hover, swipe)
@@ -0,0 +1,82 @@
# REST API Design Guide
Practical principles for designing consistent, predictable, and evolvable HTTP APIs.
## URL Naming Conventions
- Use plural nouns for collections: `/orders`, `/users`, `/products`
- Nest resources to express ownership: `/users/{userId}/orders/{orderId}`
- Keep URLs shallow (max 2-3 levels); flatten when relationships are weak
- Use kebab-case for multi-word segments: `/order-items`, not `/orderItems`
- Avoid verbs in URLs; let HTTP methods convey the action
- Use query parameters for filtering, sorting, and pagination: `/orders?status=pending&sort=-createdAt`
## HTTP Method Semantics
| Method | Purpose | Idempotent | Safe |
|--------|---------|------------|------|
| GET | Retrieve resource(s) | Yes | Yes |
| POST | Create a resource or trigger a process | No | No |
| PUT | Full replace of a resource | Yes | No |
| PATCH | Partial update of a resource | No* | No |
| DELETE | Remove a resource | Yes | No |
Use POST for actions that do not map to CRUD: `POST /orders/{id}/cancel`.
## Status Code Usage
- **200 OK** — Successful GET, PUT, PATCH, or action POST
- **201 Created** — Successful POST that created a resource; include Location header
- **204 No Content** — Successful DELETE or PUT with no response body
- **400 Bad Request** — Malformed syntax or invalid field values
- **401 Unauthorized** — Missing or invalid authentication credentials
- **403 Forbidden** — Authenticated but insufficient permissions
- **404 Not Found** — Resource does not exist
- **409 Conflict** — State conflict (duplicate, version mismatch)
- **422 Unprocessable Entity** — Valid syntax but business rule violation
- **429 Too Many Requests** — Rate limit exceeded; include Retry-After header
- **500 Internal Server Error** — Unhandled server failure
## Error Response Format
Use a consistent envelope for every error. Include a machine-readable code, a human-readable message, and optional field-level detail:
```json
{
"error": {
"code": "VALIDATION_FAILED",
"message": "One or more fields failed validation.",
"details": [
{ "field": "email", "reason": "Must be a valid email address." }
],
"requestId": "abc-123"
}
}
```
Always include a request ID for traceability.
## Pagination Patterns
- **Offset-based**: `?offset=20&limit=10` — Simple but degrades on large datasets due to OFFSET cost.
- **Cursor-based**: `?cursor=eyJpZCI6MTAwfQ&limit=10` — Encode the last-seen key as an opaque token. Preferred for DynamoDB and large datasets.
- Return pagination metadata in the response body: `nextCursor`, `hasMore`, `totalCount` (if cheap to compute).
## Versioning Strategies
- **URL path versioning** (`/v1/orders`) — Most explicit, easiest for consumers. Preferred for public APIs.
- **Header versioning** (`Accept: application/vnd.myapi.v2+json`) — Cleaner URLs but harder to discover.
- Avoid query-parameter versioning (`?version=2`); it conflates filtering with contract selection.
- Version only when you introduce breaking changes. Additive changes (new optional fields) do not require a new version.
## OpenAPI and AsyncAPI
- Maintain an OpenAPI 3.1 spec as the source of truth. Generate server stubs and client SDKs from it.
- For event-driven APIs (SNS, EventBridge, SQS), use AsyncAPI to document message schemas and channel bindings.
- Store specs in the repo alongside the code (`docs/openapi.yaml`) and validate them in CI with spectral or redocly-cli.
## HATEOAS Considerations
- Include `_links` in responses to guide clients to related actions and resources.
- Useful for complex state machines (order lifecycle) where available transitions change.
- For internal microservice APIs, HATEOAS is often unnecessary overhead; reserve it for public or partner APIs where discoverability matters.
@@ -0,0 +1,87 @@
# Code Analysis Guide
## Package & Build System Discovery
Scan the project root and common subdirectories for these markers:
| File | Build System | Language/Runtime |
|------|-------------|-----------------|
| `package.json` | npm/yarn/pnpm | JavaScript/TypeScript |
| `tsconfig.json` | TypeScript compiler | TypeScript |
| `requirements.txt` / `pyproject.toml` / `setup.py` | pip/poetry/setuptools | Python |
| `Cargo.toml` | Cargo | Rust |
| `go.mod` | Go modules | Go |
| `pom.xml` | Maven | Java/Kotlin |
| `build.gradle` / `build.gradle.kts` | Gradle | Java/Kotlin |
| `Gemfile` | Bundler | Ruby |
| `*.csproj` / `*.sln` | dotnet/MSBuild | C# |
| `Makefile` | Make | Any |
| `Dockerfile` / `docker-compose.yml` | Docker | Containerized |
| `serverless.yml` / `template.yaml` | Serverless/SAM | Cloud functions |
| `cdk.json` / `cdktf.json` | CDK/CDKTF | Infrastructure |
## Framework Detection Patterns
Identify frameworks by scanning imports and configuration:
- **React**: `import React`, `jsx`/`tsx` files, `react-dom`
- **Next.js**: `next.config.js`, `pages/` or `app/` directory structure
- **Express**: `require('express')`, `app.get/post/use` patterns
- **FastAPI**: `from fastapi import`, `@app.get` decorators
- **Django**: `settings.py` with `INSTALLED_APPS`, `urls.py`, `models.py`
- **Spring Boot**: `@SpringBootApplication`, `application.properties/yml`
- **Rails**: `config/routes.rb`, `app/controllers/`, `ActiveRecord`
## Source File Classification
Classify every source file into one of these categories:
- **Model/Entity**: Data structures, database models, DTOs, schemas
- **Controller/Handler**: Request routing, input parsing, response formatting
- **Service/UseCase**: Business logic, orchestration, domain operations
- **Repository/DAO**: Data access, queries, persistence abstraction
- **Utility/Helper**: Cross-cutting functions, formatters, validators
- **Configuration**: App config, environment setup, dependency injection
- **Middleware**: Request/response pipeline (auth, logging, error handling)
- **Test**: Unit tests, integration tests, fixtures, factories
- **Migration**: Database schema changes, data migrations
- **Static/Asset**: Templates, stylesheets, images, static content
## Dependency Graph Extraction
For each source file, extract:
1. **Direct imports** -- modules/packages this file depends on
2. **Exported symbols** -- functions/classes/constants this file provides
3. **External dependencies** -- third-party packages used
4. **Circular references** -- files that import each other (flag these)
Build a dependency adjacency list: `file -> [dependency1, dependency2, ...]`
## Code Quality Quick Assessment
Rate each of these on a 3-point scale (good/fair/poor):
- **Naming clarity**: Are variables, functions, and files self-documenting?
- **Function size**: Are functions under 30 lines with single responsibility?
- **Error handling**: Are errors caught, logged, and propagated appropriately?
- **Test presence**: Do critical paths have corresponding test files?
- **Duplication**: Are there copy-paste patterns that should be abstracted?
- **Dead code**: Are there unused imports, unreachable branches, commented-out blocks?
## API Endpoint Inventory
For each discovered endpoint, record:
- HTTP method and path (or GraphQL operation name)
- Request parameters (path, query, body, headers)
- Response shape and status codes
- Authentication/authorization requirements
- Rate limiting or throttling configuration
- Associated middleware chain
## Technical Debt Indicators
Flag these patterns during code scan:
- TODO/FIXME/HACK comments (count and categorize)
- Suppressed linter warnings (`// eslint-disable`, `# noqa`, `@SuppressWarnings`)
- Hard-coded credentials, URLs, or magic numbers
- Deeply nested conditionals (>3 levels)
- God classes/files (>500 lines with multiple responsibilities)
- Missing error handling on I/O operations
- Outdated dependencies (major version behind)
@@ -0,0 +1,136 @@
# Code Generation Guide
## Implementation Pattern Selection
Choose patterns based on the problem domain:
| Pattern | When to Use | Avoid When |
|---------|-------------|------------|
| **Repository** | Abstracting data access, multiple storage backends | Single database, simple CRUD only |
| **Service Layer** | Coordinating business logic across multiple repositories | Logic fits in a single model method |
| **Factory** | Complex object creation, conditional construction logic | Simple constructor suffices |
| **Strategy** | Runtime behavior variation (e.g., payment processing, notifications) | Only one algorithm exists |
| **Observer/Event** | Decoupling side effects from core logic (email, logging, cache invalidation) | Synchronous response required from all handlers |
| **Middleware/Pipeline** | Cross-cutting concerns (auth, logging, validation, rate limiting) | Single-purpose request handling |
| **Adapter** | Wrapping external APIs/SDKs behind a stable internal interface | Internal-only code with no external dependencies |
## Framework-Specific Generation Strategies
### General Principles (All Frameworks)
1. Scan existing code for conventions before generating new code
2. Match the project's import style (named vs. default, absolute vs. relative)
3. Follow the project's directory structure conventions
4. Use the project's established error handling pattern
5. Match existing naming conventions (camelCase, snake_case, PascalCase)
### Web API Implementation Checklist
For each endpoint, generate:
- [ ] Route definition with HTTP method and path
- [ ] Request validation (path params, query params, body schema)
- [ ] Authentication/authorization middleware
- [ ] Service call with error handling
- [ ] Response serialization with correct status code
- [ ] Error response formatting (consistent error envelope)
### Database Model Checklist
For each entity, generate:
- [ ] Model/schema definition with field types and constraints
- [ ] Indexes for queried fields and foreign keys
- [ ] Timestamps (created_at, updated_at) where appropriate
- [ ] Soft delete support if specified in requirements
- [ ] Migration file for schema changes
- [ ] Seed data for development/testing if applicable
## Brownfield Modification Best Practices
When modifying existing codebases (most common scenario):
### Before Writing Code
1. **Map the change surface**: Identify all files that will be touched
2. **Trace the call chain**: Follow the execution path from entry point to persistence
3. **Check for tests**: Find existing tests that cover the area being modified
4. **Identify conventions**: Note patterns used in surrounding code
### Modification Rules
- Match the surrounding code's style exactly, even if you prefer another style
- Do not refactor unrelated code in the same change
- Preserve existing function signatures when adding optional parameters
- Add backward-compatible defaults for new configuration
- Update existing tests to cover the changed behavior
- Add new tests for new behavior
### Common Pitfalls
- Breaking existing imports by renaming or moving files
- Changing a function's return type without updating all callers
- Adding required parameters to public APIs
- Modifying shared utility functions without checking all consumers
- Forgetting to update database migrations for schema changes
## Testing Patterns
### Unit Test Structure
Follow the Arrange-Act-Assert (AAA) pattern:
```
// Arrange: Set up preconditions and inputs
// Act: Execute the unit under test
// Assert: Verify the expected outcome
```
### What to Test per Unit
| Unit Type | Test Focus |
|-----------|------------|
| Service/Use Case | Business logic correctness, edge cases, error handling |
| Controller/Handler | Request parsing, response format, status codes, auth checks |
| Repository/DAO | Query correctness (use in-memory DB or test containers) |
| Utility/Helper | Input/output mapping, boundary values, null/undefined handling |
| Middleware | Pass-through behavior, rejection conditions, header manipulation |
### Test Data Strategy
- Use factories/builders for complex objects (avoid raw JSON literals)
- Isolate test data per test (no shared mutable fixtures)
- Use meaningful test data that reflects real scenarios
- Name test variables to express their purpose (`expiredToken`, `adminUser`, `emptyCart`)
## Code Quality Standards
### Function Design
- Maximum 30 lines per function (excluding tests)
- Single responsibility: one function does one thing
- Maximum 3 parameters; use an options object for more
- Return early to avoid deep nesting (guard clauses)
- Pure functions where possible (no side effects)
### Error Handling
- Fail fast: validate inputs at function entry
- Use typed/custom errors for domain-specific failures
- Never swallow exceptions silently (at minimum, log them)
- Propagate errors with context (wrap, do not replace)
- Distinguish between recoverable errors (retry) and fatal errors (abort)
### Naming Conventions
- Functions: verb + noun (`createUser`, `validateInput`, `calculateTotal`)
- Booleans: `is`/`has`/`should` prefix (`isActive`, `hasPermission`)
- Collections: plural nouns (`users`, `orderItems`)
- Constants: UPPER_SNAKE_CASE for true constants
- Avoid abbreviations unless universally understood (`id`, `url`, `api`)
### File Organization
- One primary export per file (class, function, or component)
- Group related files by feature/domain, not by technical layer
- Keep test files adjacent to source files (or in a mirrored `__tests__` directory)
- Index files only for public API re-exports, never for internal organization
## Automation-Friendly Code Rules
### data-testid Attributes
Add `data-testid` attributes to all interactive elements to support automated testing (E2E, integration, accessibility audits):
- **Required on**: buttons, inputs, links, form elements, modals, dropdowns, tabs, and other interactive containers
- **Naming convention**: `{component}-{element-role}` (e.g., `login-form-submit-button`, `user-profile-edit-link`, `settings-modal-close`)
- **Rules**:
- Use lowercase kebab-case
- Keep `data-testid` values stable across code changes — do not tie them to dynamic state or auto-generated IDs
- Avoid dynamic or auto-generated IDs (e.g., `button-${index}`) — use semantic names instead
- Group related elements under a container `data-testid` (e.g., `user-table` wrapping `user-table-row-{id}`)
- Apply to both visible and programmatically interactive elements (e.g., hidden file inputs triggered by a button)
@@ -0,0 +1,179 @@
# Code Generation Patterns
## Purpose
Standards and patterns for generating clean, maintainable code. These principles apply regardless of language and ensure generated code is production-worthy, not prototype-quality.
## Clean Code Principles
### Naming
- **Variables**: Describe what it holds, not the type. Use `customerEmail` not `str1` or `data`.
- **Functions**: Describe what it does using a verb. Use `calculateShippingCost()` not `process()` or `doWork()`.
- **Booleans**: Use `is`, `has`, `can`, `should` prefixes. Use `isActive` not `active` or `flag`.
- **Constants**: Use SCREAMING_SNAKE_CASE for true constants. Use `MAX_RETRY_COUNT` not `num`.
- **Classes**: Use nouns. Use `OrderProcessor` not `ProcessOrders` or `OrderHelper`.
### Functions
- Do one thing. If the function name includes "and", split it into two functions.
- Keep parameter count low (0-3 ideal; more than 3 suggests a parameter object is needed).
- Avoid boolean parameters that switch behavior — use two clearly named functions instead.
- Return early for guard clauses to reduce nesting depth.
### Files
- One primary concept per file (one class, one module, one component).
- Keep files under 300 lines. If longer, the concept is likely doing too much.
- Group related files by feature, not by type (see "Code Organization" below).
## SOLID Principles Applied
### Single Responsibility (S)
Each module/class has one reason to change.
```
// Bad: UserService handles auth, profile, and notifications
// Good: AuthService, ProfileService, NotificationService
```
### Open/Closed (O)
Extend behavior without modifying existing code. Use interfaces and composition.
```
// Bad: Adding a new payment type requires modifying PaymentProcessor
// Good: PaymentProcessor accepts a PaymentStrategy interface; add new strategies without touching the processor
```
### Liskov Substitution (L)
Subtypes must be substitutable for their base types without breaking behavior. If overriding a method changes the contract, the inheritance hierarchy is wrong.
### Interface Segregation (I)
No client should depend on methods it does not use. Prefer small, focused interfaces over large ones.
```
// Bad: interface Repository { find, save, delete, export, import, backup }
// Good: interface Readable { find }, interface Writable { save, delete }
```
### Dependency Inversion (D)
High-level modules depend on abstractions, not concrete implementations. Pass dependencies in; do not construct them internally.
```
// Bad: class OrderService { db = new PostgresDB() }
// Good: class OrderService { constructor(db: Database) }
```
## Error Handling Patterns
### Result Type (Preferred)
Return success/failure explicitly instead of throwing exceptions for expected failures.
```typescript
type Result<T, E> = { ok: true; value: T } | { ok: false; error: E };
function parseEmail(input: string): Result<Email, ValidationError> {
if (!isValidEmail(input)) {
return { ok: false, error: new ValidationError('Invalid email format') };
}
return { ok: true, value: new Email(input) };
}
```
### Try-Catch Hierarchy
When using exceptions, follow this hierarchy:
1. **Catch specific exceptions first** — handle known, recoverable errors
2. **Let unknown exceptions propagate** — do not catch `Exception` or `Error` broadly
3. **Catch at boundaries** — API handlers, event processors, and CLI entry points are appropriate catch-all locations
4. **Never swallow exceptions silently** — `catch (e) {}` hides bugs
### Error Messages
- Include what happened, why it happened, and what the user/developer can do about it
- Include relevant context (IDs, input values, operation being performed)
- Do not expose internal implementation details in user-facing errors
## Logging Standards
### Log Levels
- **ERROR**: Something failed that requires attention. Include enough context to diagnose.
- **WARN**: Something unexpected happened but was handled. May indicate a developing issue.
- **INFO**: Significant business events (order placed, user registered, payment processed). One per operation.
- **DEBUG**: Detailed diagnostic information for development. Never log sensitive data at any level.
### Structured Logging
```json
{
"level": "INFO",
"message": "Order placed successfully",
"orderId": "ord-123",
"customerId": "cust-456",
"totalAmount": 99.99,
"timestamp": "2024-01-15T10:30:00Z",
"correlationId": "req-789"
}
```
### Rules
- Log at operation boundaries (start, success, failure) — not inside loops
- Include correlation/request IDs for tracing across services
- Never log: passwords, tokens, API keys, PII, credit card numbers
- Use structured format (JSON) not free-text strings for production logs
## Input Validation at Boundaries
### Boundary Definition
Validate input at every trust boundary — where data enters your system from an untrusted source:
- API request handlers (HTTP, gRPC, GraphQL)
- Event/message consumers (SQS, Kafka, EventBridge)
- File upload processors
- CLI argument parsers
- Database query results from external systems
### Validation Strategy
```
External Input → Validate at boundary → Convert to domain type → Domain logic uses typed values
```
### What to Validate
- **Presence**: Required fields exist and are not null/empty
- **Type**: Values are the expected type (string, number, date)
- **Range**: Numbers within acceptable bounds, strings within length limits
- **Format**: Emails, URLs, dates, phone numbers match expected patterns
- **Business rules**: Values are valid in context (status transitions, referential integrity)
### After Validation
Once data passes the boundary and is converted to a domain type, internal code should NOT re-validate. Trust the boundary. This keeps domain logic clean and focused on business rules.
## Code Organization
### By Feature (Recommended)
```
/src
/orders
order.ts # domain model
order-service.ts # business logic
order-repository.ts # data access
order-handler.ts # API handler
order.test.ts # tests
/customers
customer.ts
customer-service.ts
...
```
### By Layer (Avoid for Medium-Large Projects)
```
/src
/models # all domain models from all features
/services # all services from all features
/repositories # all repositories from all features
/handlers # all handlers from all features
```
### Why Feature-Based is Better
- Related code is co-located — changes to "orders" touch files in one directory
- Easy to understand the scope of a feature by looking at one directory
- Supports extraction to microservices later (each feature directory is a candidate)
- Layer-based organization scatters related code across the entire project
## Code Generation Checklist
- [ ] Follows naming conventions (descriptive, consistent, no abbreviations)
- [ ] Functions are small and single-purpose
- [ ] Error handling is explicit (Result type or catch at boundaries)
- [ ] Input validation at all trust boundaries
- [ ] Structured logging at operation boundaries
- [ ] No hardcoded secrets, URLs, or configuration values
- [ ] Dependencies are injected, not constructed internally
- [ ] Tests accompany all generated code
- [ ] Code is organized by feature, not by layer
@@ -0,0 +1,74 @@
# Data Modelling Patterns
Guidance for designing data models across relational and NoSQL databases, with emphasis on DynamoDB single-table design.
## Relational vs NoSQL — When to Choose
| Factor | Relational (RDS/Aurora) | NoSQL (DynamoDB) |
|--------|------------------------|------------------|
| Access patterns | Ad-hoc, complex joins | Known, predictable queries |
| Consistency | Strong ACID transactions | Eventual by default, optional strong |
| Scale model | Vertical (read replicas for reads) | Horizontal, automatic partitioning |
| Schema | Fixed, enforced at DB level | Flexible, enforced in application |
| Cost model | Instance-hours | Request-based (on-demand) or provisioned |
Choose relational when you need complex reporting, ad-hoc queries, or multi-table transactions. Choose DynamoDB when access patterns are well-defined, you need single-digit-ms latency at any scale, or you want zero operational overhead.
## Normalization Forms (Relational)
- **1NF**: Eliminate repeating groups; every column holds atomic values.
- **2NF**: Remove partial dependencies; every non-key column depends on the full primary key.
- **3NF**: Remove transitive dependencies; non-key columns depend only on the primary key, not on other non-key columns.
- **BCNF**: Every determinant is a candidate key.
- Normalize to 3NF for transactional systems. Denormalize selectively for read-heavy workloads (materialized views, read replicas).
## DynamoDB Single-Table Design
Single-table design stores multiple entity types in one table using overloaded partition and sort keys.
**Process**:
1. List all entities and their relationships (user, order, order-item, payment).
2. Document every access pattern with the query it must serve.
3. Design PK/SK patterns to satisfy those queries. Use prefixes: `PK=USER#123`, `SK=ORDER#2024-01-15#456`.
4. Use GSIs to support additional access patterns (inverted index, sparse index).
**Key Design Rules**:
- Partition key should distribute load evenly; avoid hot partitions.
- Sort key enables range queries and hierarchical data (`SK BEGINS_WITH 'ORDER#'`).
- Use composite sort keys for multi-level queries: `SK=STATUS#pending#DATE#2024-01-15`.
- Store item collections (1:N relationships) under the same partition key for transactional writes.
## Entity-Relationship Modelling
- Start with a conceptual ER diagram: entities, attributes, relationships (1:1, 1:N, M:N).
- For M:N relationships in DynamoDB, use an adjacency list pattern: the relationship itself is an item with `PK=ENTITY_A#id`, `SK=ENTITY_B#id`.
- For relational databases, use a join table with foreign keys to both sides.
- Document cardinality and optionality; they drive schema decisions.
## Index Design
**DynamoDB GSI/LSI**:
- GSI (Global Secondary Index): Different partition key; eventually consistent. Use for alternate query patterns.
- LSI (Local Secondary Index): Same partition key, different sort key; supports strong consistency. Must be defined at table creation.
- Limit GSIs to 5-8; each consumes additional write capacity.
- Use sparse indexes (only items with the indexed attribute appear) for filtered queries.
**Relational B-tree Indexes**:
- Index columns that appear in WHERE, JOIN, and ORDER BY clauses.
- Use composite indexes for multi-column queries; put the most selective column first.
- Avoid over-indexing; each index slows writes and consumes storage.
- Use EXPLAIN to validate query plans use the intended index.
## Schema Versioning
- Add a `schemaVersion` attribute to every item (DynamoDB) or a `schema_version` column (relational).
- Application code handles backward-compatible reads across versions.
- For relational, use migration tools (Flyway, Liquibase) with numbered, idempotent scripts.
- Never drop columns in production without a deprecation period; add new columns as nullable first.
## Data Migration Strategies
- **Dual-write**: Write to old and new stores simultaneously during transition. Complex but zero-downtime.
- **ETL batch migration**: Export, transform, load. Use for one-time moves with a maintenance window.
- **Change data capture (CDC)**: Stream changes from source to target (DynamoDB Streams, RDS event notifications). Best for live migrations.
- Always run migration with a dry-run/validation pass before committing. Compare row counts and checksums.
@@ -0,0 +1,122 @@
# Reverse Engineering Artifact Templates
## Output Structure
All RE artifacts are created under `aidlc/spaces/<active-space>/codekb/<repo>/` — the durable per-repo code knowledge base shared across intents (the space-level directory the `codekb-path --repo <repo>` tool resolves).
### Required Artifacts
1. **business-overview.md** — Business domain context, purpose, key functionality
2. **architecture.md** — System architecture, patterns, component relationships, Mermaid diagrams
3. **code-structure.md** — Package/module organization, file classification, code patterns
4. **api-documentation.md** — External and internal API surfaces, endpoints, contracts
5. **component-inventory.md** — Complete component list with responsibilities and dependencies
6. **technology-stack.md** — Languages, frameworks, libraries with versions
7. **dependencies.md** — External dependencies, internal cross-package dependencies
8. **code-quality-assessment.md** — Test coverage, linting, CI/CD, documentation quality, tech debt
9. **reverse-engineering-timestamp.md** - Records when reverse engineering was performed (date, commit hash if available) plus the structured Scope of Analysis block (template below). The scope block is machine-read by `codekb-scope-diff` on the next rerun, so its accuracy decides whether a future intent can reuse the verified coverage or must merge/replace it.
### Developer Code Scan Template
```markdown
## Developer Code Scan Results
### Scan Coverage
- **Analyzed deeply**: [repo-relative dirs/files actually read and understood, one per line]
- **Skimmed only**: [areas noted at directory granularity without deep reading]
### Packages Found
- [package name] — [type] — [language] — [purpose]
### Build System
- **Type**: [build system]
- **Config Files**: [list]
- **Build Dependencies**: [package → package relationships]
### APIs Discovered
- [API type] — [location] — [endpoints/methods count]
### Frameworks & Libraries
- [name] — [version] — [purpose]
### Test Coverage
- **Test Directories**: [list]
- **Test Frameworks**: [list]
- **Coverage Config**: [present/absent]
### Code Quality Indicators
- **Linting**: [tool and config location]
- **CI/CD**: [pipeline files found]
- **Documentation**: [README presence, doc comments quality]
### Technical Debt Signals
- [signal description and location]
## Handoff Summary
- **Intent-relevant finding**: [the finding most relevant to the active intent, with file/line evidence]
- **Risks / follow-up**: [facts the architect or next stage must preserve; "None" if absent]
```
### Architecture Synthesis Template
```markdown
## Architecture Analysis
### System Overview
[High-level description of the system]
### Architectural Style
[Monolithic / Microservices / Serverless / Hybrid — with evidence]
### Component Relationships
[Mermaid diagram showing component interactions]
### Data Flow
[How data moves through the system]
### Key Design Decisions
[Notable architectural choices and their implications]
### Improvement Opportunities
[Areas where the architecture could be strengthened]
```
### Scope of Analysis Block (reverse-engineering-timestamp.md)
End reverse-engineering-timestamp.md with exactly this fenced block, filled
honestly from the Scan Coverage the developer reported - record what the run
ACTUALLY covered deeply, not what the stage aspired to cover:
````markdown
## Scope of Analysis
```yaml
scope_version: 1
kind: partial
intent: [active intent slug]
fingerprint: [output of the mint command in stage Step 3 - verbatim; it prints "unknown" when not computable]
analyzed:
paths:
- [repo-relative dir (trailing slash) or file analyzed deeply, one per line]
components:
- [component names exactly as they appear in component-inventory.md]
shallow:
paths:
- [areas only skimmed]
```
````
Rules:
- `kind: full` only when the scan genuinely covered the whole repo deeply; `analyzed.paths` MUST include the repo root (`./`). Anything less is `kind: partial`.
- `kind: partial` MUST NOT include `./` in `analyzed.paths`.
- `analyzed.paths` entries are repo-relative, directories end with `/`, no glob characters.
- Component names must match `component-inventory.md` headings verbatim - the rerun guard compares them literally.
- A full rescan wholesale replaces all 9 artifacts and builds this block only from the new run.
- For a focused scan of an existing store, read all 9 existing artifacts and the prior Scope of Analysis block first. Update or extend prose for the newly analyzed area and preserve prior prose outside it.
- With a CURRENT store, merge `analyzed.paths` and `analyzed.components` as the union of the store and this run. A CURRENT `kind: full` store remains full and retains `./`; otherwise the merged block is partial and cannot claim `./`.
- With a STALE or UNVERIFIED store, record only this run in `analyzed.paths` and `analyzed.components`, preserve the prior prose, and demote the store's prior analyzed paths into `shallow.paths`.
- With an UNKNOWN_SCOPE legacy store, merge the prior prose best-effort but record only this run in the new scope block.
- Mint `fingerprint` over the final `analyzed.paths` in the merged or replaced block.
- Build all 9 candidate artifacts under the temporary `<record>/.aidlc-codekb-stage-<repo>/` directory. Never write a cumulative merge directly into the shared CodeKB.
- The pre-scan `codekb-snapshot` paths bound verified coverage. If the scan discovers a deep path outside that set, take a new snapshot and repeat the scan over the expanded set.
- Publish only through `codekb-publish` with the snapshot's store generation and source fingerprint. A store-generation conflict requires re-reading and re-merging the winner's store; a source conflict requires a fresh scan.
@@ -0,0 +1,106 @@
# DevSecOps Pipeline Patterns
Integrating security into every stage of the CI/CD pipeline, shifting detection left and enforcing automated gates.
## Shift-Left Security Principles
- Find vulnerabilities as early as possible: in the IDE, at commit time, and during CI — not after deployment.
- Make security checks fast and non-blocking in early stages (warnings), then enforcing (gates) in later stages.
- Treat security findings like bugs: track, prioritize, and fix them within normal sprint cycles.
- Security is a shared responsibility, not a gatekeeping function. Enable developers with tooling and education.
## SAST — Static Application Security Testing
**Purpose**: Analyse source code for vulnerabilities without executing it.
**Tools**:
- **Amazon CodeGuru Reviewer**: ML-powered code review for Java and Python. Integrates with CodeCommit and GitHub.
- **SonarQube / SonarCloud**: Multi-language static analysis. Covers security, reliability, maintainability.
- **Semgrep**: Lightweight, pattern-based scanning. Fast, supports custom rules, good for enforcing team standards.
- **Bandit** (Python), **ESLint security plugins** (JavaScript/TypeScript).
**Pipeline Integration**:
- Run SAST on every pull request. Report findings as PR comments or annotations.
- Define severity thresholds: block merges on Critical/High findings; warn on Medium.
- Maintain a suppression file for accepted risks (with justification and expiry date).
## DAST — Dynamic Application Security Testing
**Purpose**: Test the running application by sending crafted requests to find runtime vulnerabilities.
- Run against staging or ephemeral environments, not production.
- Tools: OWASP ZAP (open source), Burp Suite (commercial), AWS Inspector for network scanning.
- Integrate into the CD pipeline after deployment to a test environment.
- DAST complements SAST; each catches different vulnerability classes.
## Dependency Vulnerability Scanning
**Purpose**: Detect known CVEs in third-party libraries and transitive dependencies.
**Tools**:
- **Amazon Inspector**: Scans EC2 instances, Lambda functions, and ECR images for software vulnerabilities.
- **Snyk**: Developer-focused, supports npm, pip, Maven, Go modules. Provides fix PRs.
- **Dependabot** (GitHub native): Automated dependency update PRs with vulnerability alerts.
- **npm audit / pip-audit / cargo audit**: Built-in language-level tools for quick local checks.
**Pipeline Integration**:
- Scan on every build. Fail the build on Critical/High severity CVEs with known exploits.
- Generate an SBOM (Software Bill of Materials) using Syft or Trivy for supply chain transparency.
- Review and update dependencies at least monthly; automate with Dependabot or Renovate.
## IaC Security Scanning
**Purpose**: Detect misconfigurations in infrastructure-as-code before deployment.
**Tools**:
- **cfn-lint**: CloudFormation linter. Validates syntax and best practices.
- **cfn-nag**: CloudFormation static analysis. Finds overly permissive IAM policies, unencrypted resources.
- **Checkov**: Multi-framework scanner (CloudFormation, Terraform, CDK, Kubernetes). Policy-as-code with custom rules.
- **cdk-nag**: CDK-native. Checks CDK constructs against AWS Solutions rules and NIST/HIPAA packs.
**Pipeline Integration**:
- Run IaC scanning before `cdk synth` or `cfn deploy`. Fail the pipeline on High findings.
- Use cdk-nag as a CDK Aspect so violations are caught at synthesis time, not after.
- Maintain exception rules in code (not out-of-band) with mandatory justification comments.
## Secret Detection
**Purpose**: Prevent credentials, API keys, and tokens from being committed to source control.
**Tools**:
- **git-secrets** (AWS Labs): Pre-commit hook that blocks patterns matching AWS credentials.
- **truffleHog**: Scans git history for high-entropy strings and known secret patterns.
- **Gitleaks**: Fast, configurable, supports CI and pre-commit. Good default ruleset.
- **GitHub secret scanning**: Built-in for GitHub repos; alerts on committed secrets from known providers.
**Best Practices**:
- Install pre-commit hooks for secret detection on every developer machine.
- Run secret scanning in CI as a backup for missed pre-commit hooks.
- If a secret is committed: revoke immediately, rotate, then clean git history.
- Store secrets in AWS Secrets Manager or SSM Parameter Store, never in code or environment files.
## Container Image Scanning
**Purpose**: Detect OS-level and application-level vulnerabilities in container images.
- **Amazon ECR image scanning**: Basic (Clair-based) and Enhanced (Inspector-powered) scanning.
- **Trivy**: Comprehensive open-source scanner. Covers OS packages, language libraries, and IaC files.
- Scan images on push to ECR and before deployment. Block deployment of images with Critical findings.
- Use minimal base images (distroless, Alpine) to reduce attack surface.
- Rebuild and rescan images weekly to pick up newly disclosed CVEs.
## Security Gates in CI/CD
Define clear pass/fail criteria at each pipeline stage:
| Stage | Gate | Action on Failure |
|-------|------|-------------------|
| Commit | Secret detection | Block commit (pre-commit hook) |
| PR | SAST scan | Block merge on Critical/High |
| Build | Dependency scan | Fail build on Critical with exploit |
| Build | IaC scan | Fail build on High |
| Deploy to staging | DAST scan | Block promotion to production |
| Deploy to prod | Image scan | Block deployment on Critical |
| Post-deploy | Inspector continuous scan | Alert and create ticket |
Automate exceptions with time-boxed waivers that require security team approval and auto-expire.
@@ -0,0 +1,188 @@
# NFR Requirements Guide
## Performance Requirement Benchmarking
### How to Define Performance Requirements
Every performance requirement must specify:
- **Metric**: What is being measured (response time, throughput, resource usage)
- **Target**: The quantitative threshold
- **Percentile**: Which percentile the target applies to (p50, p95, p99)
- **Load condition**: Under what concurrency/traffic the target must hold
- **Measurement method**: How and where the metric is captured
### Common Performance Targets by Application Type
| Application Type | Metric | Target | Percentile | Notes |
|-----------------|--------|--------|------------|-------|
| Web application (interactive) | Page load time | < 2s | p95 | First contentful paint |
| REST API (synchronous) | Response time | < 200ms | p99 | Excludes network transit |
| Search query | Result delivery | < 500ms | p95 | Including ranking |
| Batch processing | Throughput | N records/min | Average | Define minimum acceptable rate |
| Real-time messaging | End-to-end latency | < 100ms | p99 | From send to delivery |
| File upload | Processing time | < 5s per MB | p95 | Post-upload processing |
### Performance Anti-Requirements
Explicitly exclude unreasonable expectations:
- "The system should be fast" (not measurable)
- "All pages load instantly" (no target, no percentile)
- "Handle unlimited users" (no capacity boundary)
Replace with: "The dashboard page loads in < 3 seconds at p95 under 500 concurrent users."
## Scalability Assessment Frameworks
### Capacity Planning Template
| Dimension | Current | 6-Month Target | 12-Month Target | Scaling Mechanism |
|-----------|---------|---------------|-----------------|-------------------|
| Concurrent users | N | N * X | N * Y | Horizontal auto-scaling |
| Data volume | N GB | N * X GB | N * Y GB | Partitioning strategy |
| Requests per second | N | N * X | N * Y | Load balancer + replicas |
| Storage growth rate | N GB/month | N * X GB/month | N * Y GB/month | Tiered storage policy |
### Scalability Requirement Format
```
NFR-SCALE-NNN: [Scalability Requirement]
Current baseline: [measured current capacity]
Target capacity: [required capacity at specific timeline]
Growth model: [linear, exponential, seasonal, step-function]
Scaling approach: [horizontal, vertical, partitioning, caching]
Cost constraint: [infrastructure cost must remain below $X/month at target scale]
Degradation policy: [what degrades first when approaching capacity limits]
```
### Scaling Decision Matrix
| Signal | Scale Up | Scale Out | Optimize First |
|--------|----------|-----------|----------------|
| CPU consistently > 70% | Single-instance workloads | Stateless services | Check for inefficient algorithms |
| Memory pressure | In-memory processing | Distributed cache | Check for memory leaks |
| I/O wait > 30% | Faster storage tier | Read replicas, sharding | Query optimization, indexing |
| Queue depth growing | Faster consumers | More consumer instances | Batch size tuning |
| Connection pool exhausted | Larger pool size | Connection multiplexing | Connection leak investigation |
## Reliability Target-Setting
### SLA/SLO/SLI Hierarchy
| Term | Definition | Example | Who Defines |
|------|-----------|---------|-------------|
| **SLI** (Indicator) | The metric being measured | Successful requests / total requests | Engineering |
| **SLO** (Objective) | Internal target for the SLI | 99.95% success rate over 30 days | Engineering + Product |
| **SLA** (Agreement) | External contractual commitment | 99.9% uptime with financial penalties | Business + Legal |
**Rule**: SLO must be stricter than SLA. If SLA is 99.9%, set SLO at 99.95% to provide an internal buffer.
### Availability Targets and Their Implications
| Availability | Downtime / Year | Downtime / Month | Requires |
|-------------|-----------------|-------------------|----------|
| 99% (two nines) | 3.65 days | 7.3 hours | Basic monitoring, manual recovery |
| 99.9% (three nines) | 8.76 hours | 43.8 minutes | Auto-restart, health checks, alerting |
| 99.95% | 4.38 hours | 21.9 minutes | Multi-AZ, automated failover, load balancing |
| 99.99% (four nines) | 52.6 minutes | 4.38 minutes | Multi-region active-active, zero-downtime deploys |
| 99.999% (five nines) | 5.26 minutes | 26.3 seconds | Fully automated everything, extensive redundancy |
### Recovery Objectives
| Objective | Definition | How to Set |
|-----------|-----------|------------|
| **RTO** (Recovery Time Objective) | Maximum acceptable time to restore service | Based on business impact per hour of downtime |
| **RPO** (Recovery Point Objective) | Maximum acceptable data loss window | Based on cost of recreating or losing data |
| **MTTR** (Mean Time to Recovery) | Average time to restore from failure | Measured operationally; target should be < RTO |
| **MTBF** (Mean Time Between Failures) | Average time between failures | Measured operationally; drives reliability investment |
## Observability Requirements
### Three Pillars Specification
#### Metrics Requirements
Define for each component:
- **Business metrics**: Conversion rate, revenue processed, active users
- **Application metrics**: Request rate, error rate, latency (RED method)
- **Infrastructure metrics**: CPU, memory, disk, network (USE method)
- **Retention**: How long to keep each metric tier (1 min granularity for 7 days, 5 min for 30 days, 1 hour for 1 year)
#### Logging Requirements
| Log Level | When to Use | Retention | Indexing |
|-----------|------------|-----------|---------|
| ERROR | Unrecoverable failures requiring attention | 90 days | Full-text indexed |
| WARN | Recoverable issues, degraded behavior | 30 days | Full-text indexed |
| INFO | Significant business events, state transitions | 30 days | Structured fields only |
| DEBUG | Diagnostic detail for troubleshooting | 7 days | Not indexed (stored only) |
#### Tracing Requirements
- Distributed traces across all service boundaries
- Trace context propagation via W3C Trace Context headers
- Sampling strategy: 100% for errors, 10% for normal traffic (adjust for volume)
- Trace retention: 7 days at full detail, 30 days for trace metadata
### Alerting Requirements Template
For each alert:
```
Alert: [descriptive name]
SLI: [which service level indicator]
Threshold: [when to fire -- e.g., error rate > 1% for 5 minutes]
Severity: [page (wake someone up) | ticket (next business day) | log (informational)]
Runbook: [link to response procedure]
Notification: [who gets notified via which channel]
Auto-remediation: [if any automated response is triggered]
```
### Observability Anti-Patterns to Avoid
- Alerting on causes instead of symptoms (alert on error rate, not CPU usage)
- Missing correlation IDs across service boundaries
- Logging sensitive data (PII, credentials, tokens)
- Alert fatigue from noisy thresholds (tune before deploying)
- Dashboard sprawl without clear ownership (every dashboard needs an owner)
## Security Requirement Templates
### Authentication Requirements
```
NFR-AUTH-NNN: [Authentication Requirement]
Method: [OAuth 2.0, JWT, session-based, API key, mTLS]
Token lifetime: [access token TTL, refresh token TTL]
MFA requirement: [none, optional, required for admin, required for all]
Session management: [max concurrent sessions, idle timeout, absolute timeout]
Password policy: [min length, complexity, rotation, breach detection]
```
### Authorization Requirements
```
NFR-AUTHZ-NNN: [Authorization Requirement]
Model: [RBAC, ABAC, ACL, policy-based]
Roles: [list of roles and their permission boundaries]
Resource granularity: [organization, team, user, resource-level]
Delegation: [can users delegate permissions, under what constraints]
Audit: [which authorization decisions are logged]
```
### Data Protection Requirements
```
NFR-DATA-NNN: [Data Protection Requirement]
Classification: [public, internal, confidential, restricted]
Encryption at rest: [algorithm, key management, rotation schedule]
Encryption in transit: [TLS version, cipher suites, certificate management]
PII handling: [identification, masking, pseudonymization, retention limits]
Data residency: [geographic constraints, cross-border transfer rules]
Backup encryption: [separate key, recovery testing frequency]
```
### Compliance Requirements
```
NFR-COMP-NNN: [Compliance Requirement]
Framework: [GDPR, SOC 2, HIPAA, PCI-DSS, ISO 27001]
Scope: [which components/data are in scope]
Controls: [specific controls required — access logging, encryption, retention]
Audit frequency: [continuous, quarterly, annual]
Evidence requirements: [what documentation/logs must be maintained]
```
### Security Anti-Requirements
Explicitly exclude unreasonable expectations:
- "The system should be secure" (not measurable)
- "No vulnerabilities" (impossible — define acceptable risk)
- "Military-grade encryption" (undefined — specify algorithm and key length)
Replace with: "All API endpoints require JWT authentication with RS256 signing. Access tokens expire after 15 minutes. Refresh tokens expire after 7 days and are single-use."
@@ -0,0 +1,77 @@
# Security Guide
## OWASP Top 10 (Application Security Checklist)
For every application, verify defenses against each category:
1. **Broken Access Control**: Enforce authorization on every endpoint. Deny by default. Verify object-level access (IDOR prevention). Disable directory listing. Invalidate sessions on logout.
2. **Cryptographic Failures**: Use TLS 1.2+ for all data in transit. Encrypt sensitive data at rest (AES-256). Never store passwords in plaintext (use bcrypt/argon2 with salt). Do not roll custom crypto. Classify data sensitivity and protect accordingly.
3. **Injection**: Parameterize all database queries (no string concatenation). Use ORM methods for queries. Validate and sanitize all input. Apply Content Security Policy headers. Encode output contextually (HTML, URL, JavaScript, CSS).
4. **Insecure Design**: Threat model during design, not after. Limit resource consumption per user. Use secure design patterns (see below). Separate business logic from security controls.
5. **Security Misconfiguration**: Remove default accounts and passwords. Disable unnecessary features and services. Set security headers (HSTS, CSP, X-Frame-Options, X-Content-Type-Options). Keep dependencies patched. Review cloud service configurations (S3 bucket policies, security groups).
6. **Vulnerable Components**: Maintain a software bill of materials (SBOM). Scan dependencies weekly. Pin dependency versions. Have a process for emergency patching of critical CVEs. Remove unused dependencies.
7. **Authentication Failures**: Implement rate limiting on login (5 attempts per minute). Enforce password complexity (12+ characters, no common passwords). Support MFA. Use secure session management (HttpOnly, Secure, SameSite cookies). Implement account lockout with notification.
8. **Data Integrity Failures**: Verify integrity of software updates and CI/CD pipelines. Use signed artifacts. Validate data from untrusted sources. Protect deserialization (avoid accepting serialized objects from users).
9. **Logging & Monitoring Failures**: Log authentication events (success and failure). Log access control failures. Log input validation failures. Do NOT log sensitive data (passwords, tokens, PII). Ship logs to centralized, tamper-resistant storage. Set up alerts for suspicious patterns.
10. **SSRF (Server-Side Request Forgery)**: Validate and whitelist URLs for server-side requests. Block requests to internal networks (169.254.x.x, 10.x.x.x, 172.16-31.x.x). Disable HTTP redirects for server-initiated requests. Do not expose raw error messages from server-side requests.
## STRIDE Threat Modeling
For each component and data flow, assess:
| Threat | Question | Example Mitigation |
|--------|----------|-------------------|
| **S**poofing | Can an attacker impersonate a user or service? | Authentication, mutual TLS, API keys |
| **T**ampering | Can data be modified in transit or at rest? | Input validation, checksums, signed tokens |
| **R**epudiation | Can a user deny performing an action? | Audit logging, non-repudiation controls |
| **I**nformation Disclosure | Can sensitive data leak? | Encryption, access controls, data masking |
| **D**enial of Service | Can the system be made unavailable? | Rate limiting, autoscaling, circuit breakers |
| **E**levation of Privilege | Can a user gain unauthorized permissions? | Least privilege, RBAC enforcement, input validation |
## Authentication & Authorization Patterns
### Authentication
- **Session-based**: Server-side sessions with HttpOnly/Secure/SameSite cookies. Best for server-rendered web apps.
- **JWT**: Stateless tokens with short expiry (15 min access, 7 day refresh). Best for SPAs and APIs. Store access token in memory, refresh token in HttpOnly cookie.
- **API Keys**: For service-to-service communication. Rotate regularly. Scope to minimum permissions.
- **OAuth2/OIDC**: For third-party authentication delegation. Use authorization code flow with PKCE. Never use implicit flow.
### Authorization
- **RBAC (Role-Based)**: Assign permissions to roles, roles to users. Good for well-defined hierarchies.
- **ABAC (Attribute-Based)**: Evaluate rules based on user, resource, action, and environment attributes. Good for complex, context-dependent policies.
- **Object-Level**: Always verify the requesting user has access to the specific resource being requested. Never trust client-provided ownership claims.
## Data Protection Requirements
Classify data into tiers and apply controls:
| Tier | Examples | At Rest | In Transit | Access | Retention |
|------|----------|---------|------------|--------|-----------|
| Public | Marketing content | None required | HTTPS preferred | Open | Indefinite |
| Internal | Business docs | Encrypted volume | HTTPS required | Authenticated | Per policy |
| Confidential | PII, financial | AES-256, key rotation | TLS 1.2+ required | Role-restricted | Minimized |
| Restricted | Passwords, keys | HSM/KMS, separate storage | mTLS | Named individuals | Shortest possible |
## Secure Coding Practices Checklist
For code review, verify:
- [ ] All user input validated (type, length, range, format)
- [ ] SQL queries parameterized (no string interpolation)
- [ ] Output encoded for context (HTML, URL, JS)
- [ ] Authentication checked on every protected endpoint
- [ ] Authorization checked for the specific resource being accessed
- [ ] Sensitive data not logged (passwords, tokens, PII)
- [ ] Error messages do not reveal internal details (stack traces, SQL errors)
- [ ] File uploads validated (type, size, scanned for malware)
- [ ] CORS configured to allow only expected origins
- [ ] Rate limiting applied to authentication and expensive operations
- [ ] Secrets loaded from environment/secrets manager, never hardcoded
@@ -0,0 +1,100 @@
# Threat Modelling with STRIDE
A structured approach to identifying, classifying, and mitigating security threats during the design phase.
## STRIDE Categories
| Category | Threat | Violated Property | Example |
|----------|--------|-------------------|---------|
| **S — Spoofing** | Attacker pretends to be another user or system | Authentication | Stolen JWT used to access another user's data |
| **T — Tampering** | Attacker modifies data in transit or at rest | Integrity | Man-in-the-middle alters API request payload |
| **R — Repudiation** | Attacker denies performing an action | Non-repudiation | User disputes a financial transaction with no audit trail |
| **I — Information Disclosure** | Sensitive data exposed to unauthorized parties | Confidentiality | Error message leaks stack trace and database schema |
| **D — Denial of Service** | System made unavailable to legitimate users | Availability | Unbounded API request floods Lambda concurrency |
| **E — Elevation of Privilege** | Attacker gains higher access than authorized | Authorization | Regular user exploits IDOR to access admin endpoints |
## Threat Modelling Process
### Step 1: Define Scope
Identify what you are threat-modelling: a single microservice, a complete feature, or an entire system. Smaller scopes produce more actionable results.
### Step 2: Create a Data Flow Diagram (DFD)
Draw the system showing:
- **External entities**: Users, third-party systems, partner APIs
- **Processes**: Lambda functions, ECS services, API Gateway
- **Data stores**: DynamoDB tables, S3 buckets, RDS instances
- **Data flows**: Arrows showing data movement, labelled with protocol and data type
- **Trust boundaries**: Lines separating zones of different trust levels (public internet, VPC, private subnet)
### Step 3: Identify Threats
Walk through each element and data flow in the DFD. For each, ask the six STRIDE questions:
- Can an attacker spoof this identity?
- Can an attacker tamper with this data?
- Can an attacker deny this action?
- Can this component leak information?
- Can this component be denied service?
- Can an attacker elevate privilege through this component?
### Step 4: Assess Risk
For each identified threat, score it using likelihood and impact.
### Step 5: Define Mitigations
Map each threat to a specific countermeasure. Track mitigations as actionable tasks in the backlog.
## Risk Scoring: Likelihood x Impact
Use a simple 3x3 or 5x5 matrix:
| | Low Impact | Medium Impact | High Impact |
|---|-----------|---------------|-------------|
| **High Likelihood** | Medium | High | Critical |
| **Medium Likelihood** | Low | Medium | High |
| **Low Likelihood** | Low | Low | Medium |
- **Likelihood factors**: Attack complexity, required access level, availability of exploits, attacker motivation.
- **Impact factors**: Data sensitivity, financial loss, regulatory penalties, reputational damage, blast radius.
## The DREAD Model (Alternative Scoring)
Score each threat 1-10 on five dimensions, then average:
- **Damage**: How severe is the impact?
- **Reproducibility**: How reliably can the attack be repeated?
- **Exploitability**: How much skill/effort is required?
- **Affected Users**: How many users are impacted?
- **Discoverability**: How easy is the vulnerability to find?
DREAD is useful when STRIDE identifies many threats and you need to prioritize.
## Attack Surface Analysis
Enumerate all entry points an attacker could target:
- Public API endpoints (API Gateway, ALB)
- Authentication flows (Cognito hosted UI, custom login)
- File upload endpoints (S3 pre-signed URLs)
- WebSocket connections
- Third-party integrations (webhooks, OAuth callbacks)
- Administrative interfaces (console access, SSH, bastion)
- CI/CD pipelines (build scripts, deployment credentials)
Minimize attack surface: disable unused endpoints, restrict network access, apply least-privilege IAM.
## Mitigation Mapping
For each STRIDE category, common AWS mitigations include:
| Category | Mitigations |
|----------|-------------|
| Spoofing | Cognito + MFA, mutual TLS, API key rotation, IAM roles (no long-lived credentials) |
| Tampering | HTTPS everywhere, S3 Object Lock, DynamoDB encryption, request signing |
| Repudiation | CloudTrail logging, application audit logs, DynamoDB Streams for change history |
| Information Disclosure | Encryption at rest (KMS), VPC endpoints, security groups, suppress verbose errors |
| Denial of Service | WAF rate limiting, API Gateway throttling, Lambda reserved concurrency, Shield Advanced |
| Elevation of Privilege | Least-privilege IAM, ABAC policies, input validation, IDOR checks in application logic |
## When to Threat Model
- During design (Stage 3 of AI-DLC) before implementation begins.
- When adding a new external integration or data flow.
- When changing authentication or authorization mechanisms.
- After a security incident, to update the model with newly discovered threats.
- Review and refresh threat models at least annually.
@@ -0,0 +1,107 @@
# Incident Response Guide
Structured processes for detecting, responding to, and learning from production incidents.
## Incident Severity Levels
| Level | Name | Criteria | Response Time | Examples |
|-------|------|----------|---------------|---------|
| **SEV1** | Critical | Complete service outage or data loss affecting all users | < 15 minutes | Production down, data breach, payment processing failure |
| **SEV2** | Major | Significant degradation affecting many users, no workaround | < 30 minutes | Partial outage, error rate > 10%, major feature broken |
| **SEV3** | Minor | Limited impact, workaround available | < 2 hours | Non-critical feature broken, intermittent errors, slow performance |
| **SEV4** | Low | Cosmetic or minor issue, no user impact | Next business day | UI glitch, non-critical log errors, minor config drift |
## Escalation Matrix
Define who to contact at each severity level:
| Severity | Primary Responder | Escalation (30 min) | Escalation (1 hour) |
|----------|------------------|---------------------|---------------------|
| SEV1 | On-call engineer | Engineering manager + Incident commander | VP Engineering + Stakeholder communication |
| SEV2 | On-call engineer | Team lead | Engineering manager |
| SEV3 | On-call engineer | Team lead (if unresolved in 4 hours) | — |
| SEV4 | Any team member | — | — |
## On-Call Rotation
- Rotate weekly among team members. Ensure at least 2 people are trained for on-call at all times.
- Primary and secondary on-call: secondary takes over if primary is unreachable within 10 minutes.
- On-call handoff includes: review of active alerts, ongoing issues, recent deployments, and known risks.
- Compensate on-call fairly: time off in lieu, on-call stipend, or both.
- Maximum on-call frequency: no more than 1 week in 4. If the team is too small, address staffing.
## Incident Commander Role
For SEV1 and SEV2 incidents, designate an Incident Commander (IC) who:
- **Coordinates** response efforts; does not debug directly.
- **Communicates** status updates to stakeholders at regular intervals (every 15-30 minutes).
- **Delegates** workstreams: investigation, mitigation, communication, documentation.
- **Decides** when to escalate, when to roll back, when to declare resolution.
- **Documents** timeline, actions taken, and decisions made in a shared incident channel.
The IC is not necessarily the most senior engineer; it is the person who can coordinate effectively under pressure.
## Communication During Incidents
### Internal
- Create a dedicated Slack/Teams channel: `#incident-YYYY-MM-DD-short-description`.
- Post structured updates: **Status** (investigating/identified/mitigating/resolved), **Impact** (who is affected), **Next Step** (what we are doing), **ETA** (when the next update will be).
- Keep the channel focused; move side discussions to threads.
### External
- SEV1: Status page update within 20 minutes. Customer communication within 1 hour.
- SEV2: Status page update within 1 hour if customer-visible.
- Use pre-drafted templates for common scenarios: "We are experiencing elevated error rates..."
- Never speculate about root cause in external communications until confirmed.
## Post-Incident Review (Blameless Postmortem)
Conduct within 48 hours of incident resolution for SEV1/SEV2.
### Structure
1. **Timeline**: Minute-by-minute account from detection to resolution.
2. **Impact**: Users affected, duration, financial/data impact.
3. **Root Cause**: Technical cause(s) of the incident.
4. **Contributing Factors**: Process, tooling, or knowledge gaps that allowed the incident.
5. **What Went Well**: Effective parts of the response.
6. **What Could Be Improved**: Gaps in detection, response, or communication.
7. **Action Items**: Specific, assigned, time-boxed improvements.
### Blameless Principles
- Focus on systems and processes, not individuals.
- People made the best decisions they could with the information available.
- Ask "what" and "how", not "who".
- The goal is to improve the system so the same failure cannot recur.
## SSM Automation Runbooks
Create automated runbooks for common operational tasks and incident responses:
- **Restart service**: ECS task restart, Lambda function redeployment.
- **Scale out**: Increase desired count, adjust auto-scaling thresholds.
- **Database failover**: Trigger RDS failover, verify application reconnection.
- **Clear queue backlog**: Increase consumer concurrency, purge DLQ after inspection.
- **Rotate credentials**: Update secrets in Secrets Manager, trigger dependent service reloads.
Store runbooks as SSM Automation documents in version control. Reference them in alarm actions for automated remediation.
## Automated Remediation Patterns
- CloudWatch alarm triggers Lambda function to restart unhealthy ECS tasks.
- Auto Scaling policies respond to custom metrics (queue depth, error rate) not just CPU.
- EventBridge rules detect specific error patterns and invoke Step Functions for multi-step remediation.
- Always include circuit breakers: limit automated remediation attempts (max 3 restarts in 10 minutes) to prevent remediation loops.
## RTO and RPO Targets
- **RTO (Recovery Time Objective)**: Maximum acceptable downtime. How quickly must the system be restored?
- **RPO (Recovery Point Objective)**: Maximum acceptable data loss. How much data can we afford to lose?
| Tier | RTO | RPO | Strategy |
|------|-----|-----|----------|
| Critical (payments, auth) | < 5 minutes | 0 (zero data loss) | Multi-AZ active-active, synchronous replication |
| High (order processing) | < 30 minutes | < 5 minutes | Multi-AZ with automated failover, point-in-time recovery |
| Standard (reporting, analytics) | < 4 hours | < 1 hour | Regular backups, automated restoration |
| Low (internal tools) | < 24 hours | < 24 hours | Daily backups, manual restoration |
Test RTO/RPO targets quarterly through game days and disaster recovery drills.
@@ -0,0 +1,64 @@
# NFR Performance and Scalability Guide
> This guide supplements the full NFR Requirements Guide held by the Security Engineer (lead agent for the NFR Requirements stage). It provides the DevOps Engineer with the performance and scalability sections needed for infrastructure-focused contributions during NFR Requirements and NFR Design stages.
## Performance Requirement Benchmarking
### How to Define Performance Requirements
Every performance requirement must specify:
- **Metric**: What is being measured (response time, throughput, resource usage)
- **Target**: The quantitative threshold
- **Percentile**: Which percentile the target applies to (p50, p95, p99)
- **Load condition**: Under what concurrency/traffic the target must hold
- **Measurement method**: How and where the metric is captured
### Common Performance Targets by Application Type
| Application Type | Metric | Target | Percentile | Notes |
|-----------------|--------|--------|------------|-------|
| Web application (interactive) | Page load time | < 2s | p95 | First contentful paint |
| REST API (synchronous) | Response time | < 200ms | p99 | Excludes network transit |
| Search query | Result delivery | < 500ms | p95 | Including ranking |
| Batch processing | Throughput | N records/min | Average | Define minimum acceptable rate |
| Real-time messaging | End-to-end latency | < 100ms | p99 | From send to delivery |
| File upload | Processing time | < 5s per MB | p95 | Post-upload processing |
### Performance Anti-Requirements
Explicitly exclude unreasonable expectations:
- "The system should be fast" (not measurable)
- "All pages load instantly" (no target, no percentile)
- "Handle unlimited users" (no capacity boundary)
Replace with: "The dashboard page loads in < 3 seconds at p95 under 500 concurrent users."
## Scalability Assessment Frameworks
### Capacity Planning Template
| Dimension | Current | 6-Month Target | 12-Month Target | Scaling Mechanism |
|-----------|---------|---------------|-----------------|-------------------|
| Concurrent users | N | N * X | N * Y | Horizontal auto-scaling |
| Data volume | N GB | N * X GB | N * Y GB | Partitioning strategy |
| Requests per second | N | N * X | N * Y | Load balancer + replicas |
| Storage growth rate | N GB/month | N * X GB/month | N * Y GB/month | Tiered storage policy |
### Scalability Requirement Format
```
NFR-SCALE-NNN: [Scalability Requirement]
Current baseline: [measured current capacity]
Target capacity: [required capacity at specific timeline]
Growth model: [linear, exponential, seasonal, step-function]
Scaling approach: [horizontal, vertical, partitioning, caching]
Cost constraint: [infrastructure cost must remain below $X/month at target scale]
Degradation policy: [what degrades first when approaching capacity limits]
```
### Scaling Decision Matrix
| Signal | Scale Up | Scale Out | Optimize First |
|--------|----------|-----------|----------------|
| CPU consistently > 70% | Single-instance workloads | Stateless services | Check for inefficient algorithms |
| Memory pressure | In-memory processing | Distributed cache | Check for memory leaks |
| I/O wait > 30% | Faster storage tier | Read replicas, sharding | Query optimization, indexing |
| Queue depth growing | Faster consumers | More consumer instances | Batch size tuning |
| Connection pool exhausted | Larger pool size | Connection multiplexing | Connection leak investigation |
@@ -0,0 +1,100 @@
# Observability Patterns
Building visibility into system behaviour through metrics, logs, and traces to enable fast diagnosis and proactive detection.
## The Three Pillars of Observability
### 1. Metrics
Numeric measurements aggregated over time. Low cardinality, low cost, ideal for alerting and dashboards.
- **Types**: Counters (always increase), Gauges (can go up/down), Histograms (distribution of values).
- **AWS**: CloudWatch Metrics, CloudWatch Embedded Metric Format (EMF) for custom metrics from Lambda/ECS.
- Emit custom business metrics: orders placed per minute, payment failures per hour, sign-ups per day.
### 2. Logs
Timestamped records of discrete events. High cardinality, high volume.
- Use structured logging (JSON) with consistent fields: `timestamp`, `level`, `requestId`, `service`, `message`.
- Correlate logs across services using a shared `requestId` or `traceId` propagated through headers.
- **AWS**: CloudWatch Logs, with Log Groups per service and environment.
### 3. Traces
End-to-end path of a request through multiple services. Shows latency breakdown and dependency relationships.
- **AWS**: X-Ray for distributed tracing. Auto-instruments SDK calls to AWS services.
- Instrument custom segments for business logic, database queries, and external HTTP calls.
- Use trace maps to visualize service dependencies and identify latency bottlenecks.
## The Four Golden Signals (Google SRE)
Monitor these four signals for every service:
1. **Latency**: Time to serve a request. Measure as percentiles (p50, p95, p99). Track separately for successful and failed requests.
2. **Traffic**: Demand on the system. Requests per second for APIs, messages per second for queues, active connections for WebSocket.
3. **Errors**: Rate of failed requests. Include both explicit errors (5xx) and implicit errors (wrong responses, timeouts, retries).
4. **Saturation**: How full the system is. CPU utilization, memory usage, disk I/O, connection pool usage, Lambda concurrent executions vs reserved concurrency.
## CloudWatch Dashboards and Alarms
### Dashboard Design
- One dashboard per service with golden signals at the top.
- Include dependency health: downstream service latency, database connections, queue depth.
- Use CloudWatch Metrics Math for derived metrics: error rate = errors / (errors + successes) * 100.
- Standardize dashboard layouts across services for consistency.
### Alarm Configuration
- Alarm on symptoms, not causes. Alert on "error rate > 1%" not "CPU > 80%".
- Use composite alarms to reduce noise: only alert when multiple conditions are true simultaneously.
- Set appropriate evaluation periods: avoid single-datapoint alarms that fire on transient spikes. Use `3 out of 5 datapoints` for stability.
- Define alarm actions: SNS notification, Lambda remediation, OpsCenter item creation.
- Severity tiers: P1 (page on-call), P2 (Slack alert, fix within hours), P3 (ticket, fix within sprint).
## X-Ray Distributed Tracing
- Enable X-Ray tracing on API Gateway, Lambda, ECS, and SQS.
- Use the X-Ray SDK to create custom subsegments for application logic.
- Add annotations (indexed, searchable) for key dimensions: `customerId`, `orderStatus`, `region`.
- Add metadata (not indexed) for debugging detail: request/response bodies, query parameters.
- Use X-Ray groups and filter expressions to find traces matching specific criteria: `service("payment-service") AND responsetime > 3`.
## CloudWatch Logs Insights Queries
Useful query patterns for operational troubleshooting:
```
# Error rate over time
filter @message like /ERROR/
| stats count() as errors by bin(5m)
# Slowest requests
filter @duration > 1000
| sort @duration desc
| limit 20
# Request count by status code
parse @message '"statusCode":*,' as statusCode
| stats count() by statusCode
# Cold start impact (Lambda)
filter @type = "REPORT"
| stats avg(@duration), max(@duration), avg(@initDuration) by bin(10m)
```
## Anomaly Detection
- CloudWatch Anomaly Detection uses ML to establish a baseline and alert on deviations.
- Enable on latency and error rate metrics that have stable patterns.
- Set the detection band width based on acceptable variance (2 standard deviations is a common starting point).
- Combine anomaly detection with static thresholds: anomaly detection catches gradual drift, static thresholds catch acute failures.
## Log Aggregation and Retention
- Route all logs to CloudWatch Logs with consistent naming: `/aws/lambda/{service}`, `/ecs/{cluster}/{service}`.
- Set retention policies per log group: 30 days for development, 90 days for staging, 1-3 years for production (or per compliance requirements).
- Archive to S3 for long-term retention and cost optimization. Use Glacier for logs older than 90 days.
- Consider CloudWatch cross-account log aggregation for multi-account setups.
## Synthetic Monitoring
- CloudWatch Synthetics canaries: run scripted checks on a schedule (every 1-5 minutes).
- Test critical user journeys: login, place order, view dashboard.
- Canaries run from AWS-managed infrastructure; detect issues before real users report them.
- Alert on canary failure with P1 severity; it means the user-facing path is broken.
- Use visual monitoring (screenshot comparison) for UI-heavy applications.
@@ -0,0 +1,99 @@
# SLO, SLI, and Error Budget Patterns
Defining and managing service level objectives using Google SRE principles adapted for AWS environments.
## Key Terminology
- **SLI (Service Level Indicator)**: A quantitative measure of a specific aspect of service performance. Example: "the proportion of requests served in under 200ms."
- **SLO (Service Level Objective)**: A target value or range for an SLI. Example: "99.9% of requests will be served in under 200ms, measured over a 30-day rolling window."
- **SLA (Service Level Agreement)**: A contractual commitment to a customer, with consequences for breach. SLAs are looser than SLOs to provide a safety margin.
- **Error Budget**: The allowed amount of unreliability. If SLO is 99.9%, the error budget is 0.1% (approximately 43 minutes of downtime per 30 days).
## SLI Definition
Choose SLIs that reflect the user's experience, not internal system metrics.
### Common SLI Types
| SLI Type | Definition | Measurement |
|----------|-----------|-------------|
| **Availability** | Proportion of successful requests | `(successful requests / total requests) * 100` |
| **Latency** | Proportion of requests faster than threshold | `(requests < 200ms / total requests) * 100` |
| **Throughput** | Requests processed per unit time | CloudWatch metric: `RequestCount` per minute |
| **Error Rate** | Proportion of requests returning errors | `(5xx responses / total responses) * 100` |
| **Freshness** | Proportion of data updated within threshold | Time since last successful sync vs target |
| **Correctness** | Proportion of responses with correct output | Validated against ground truth or invariants |
### SLI Specification Best Practices
- Measure at the point closest to the user (API Gateway, ALB) not at the application.
- Exclude health check traffic and synthetic monitoring from SLI calculations.
- Use CloudWatch Metrics Math or CloudWatch Contributor Insights for SLI computation.
- Define SLIs per critical user journey, not per API endpoint.
## SLO Target Setting
### Process
1. Measure current performance for 2-4 weeks to establish a baseline.
2. Set the SLO slightly below the observed baseline (if p99 latency is consistently 150ms, set SLO at 200ms).
3. Validate with stakeholders that the target aligns with user expectations and business requirements.
4. Start conservative; tighten SLOs as reliability improves and tooling matures.
### Guidelines
- Do not set SLOs at 100%. It is impossible to achieve and eliminates the error budget for deployments and improvements.
- Typical SLO targets: 99.9% for customer-facing services, 99.5% for internal services, 99% for batch processing.
- Use a 30-day rolling window, not calendar month, to avoid reset-day gaming.
- Document the measurement methodology alongside the target.
## Error Budgets
The error budget is the inverse of the SLO: `error budget = 1 - SLO target`.
### Error Budget Policy
Define what happens when the error budget is consumed:
| Budget Status | Implication | Action |
|--------------|-------------|--------|
| > 50% remaining | Healthy | Normal feature development velocity |
| 25-50% remaining | Caution | Increase deployment monitoring, review recent incidents |
| < 25% remaining | At risk | Slow deployments, prioritize reliability work |
| Exhausted (0%) | Frozen | Halt feature releases, all engineering effort on reliability |
Error budgets create a shared language between product and engineering: "We can afford to ship this risky feature because we have budget, or we cannot because the budget is low."
## Burn Rate Alerting
Instead of alerting when the SLO is breached (too late), alert based on the rate at which the error budget is being consumed.
- **Burn rate**: How fast the error budget is being consumed relative to the window. A burn rate of 1.0 means the budget will be exactly exhausted at the end of the window.
- **Fast burn (SEV1)**: Burn rate > 14x over 1 hour and > 14x over 5 minutes. The budget will be consumed in ~2 hours. Page immediately.
- **Medium burn (SEV2)**: Burn rate > 6x over 6 hours and > 6x over 30 minutes. The budget will be consumed in ~5 days. Alert during business hours.
- **Slow burn (SEV3)**: Burn rate > 3x over 24 hours and > 3x over 6 hours. The budget will be consumed in ~10 days. Create a ticket.
Use multi-window, multi-burn-rate alerting (Google SRE workbook approach) to balance sensitivity and specificity.
## SLO-Based Decision Making
- **Should we deploy?** Check error budget. If healthy, deploy with normal confidence. If low, add extra validation or delay.
- **Should we invest in reliability?** If the SLO is consistently met with budget to spare, invest in features. If the budget is frequently exhausted, invest in reliability.
- **Should we adopt a new dependency?** Evaluate the dependency's SLO. Your service's SLO cannot exceed its least reliable critical dependency.
- **How much testing is enough?** Test until you are confident the deployment will not consume more than a defined fraction of the error budget.
## Toil Measurement
Toil is repetitive, automatable operational work that scales with service size.
- Track time spent on toil weekly: manual deployments, alert response, config changes, scaling actions.
- Target: toil should consume no more than 50% of an operations engineer's time (Google SRE standard).
- Prioritize automation of the most time-consuming toil tasks.
- Report toil reduction as a team metric alongside SLO compliance.
## SLO Dashboards
Build a dashboard per service showing:
- Current SLO compliance (percentage) vs target, with the measurement window
- Error budget remaining (absolute and percentage)
- Burn rate trend over the last 7 and 30 days
- SLI time series (latency percentiles, error rate, availability)
- Recent incidents that consumed error budget, with links to postmortems
Use CloudWatch dashboards with Metrics Math, or export to Grafana for richer visualisation.
@@ -0,0 +1,309 @@
# Branching Strategies
A menu of common branching strategies, what they look like, when to use them, and how AIDLC's Construction worktrees map onto each. When the orchestrator dispatches aidlc-pipeline-deploy-agent at Bolt boundaries, this file is the menu the agent surveys to map a team's affirmed branching strategy onto the `aidlc-worktree` tool's flags.
> **Reading practices:** see `knowledge/aidlc-shared/rules-reading.md` for empty-template detection, semantic-topic matching, and the active-space `project.md → team.md → org.md → hardcoded defaults` fallback chain. In the runbooks below, those files live under `aidlc/spaces/<active-space>/memory/`.
>
> See also `cicd-patterns.md` § "Branch Strategies" for the higher-level CI-flow context.
---
## Trunk-Based Development (default)
```
main ────────────────────────────────────────►
▲ ▲ ▲ ▲ ▲ ▲
│ │ │ │ │ │ short-lived feature branches
│ │ │ │ │ │ (1-2 days max), squash-merge to main
bolt-1 bolt-3 bolt-5
bolt-2 bolt-4
```
**Shape.** All work merges to `main` via short-lived feature branches. Long-lived branches don't exist. Feature flags gate incomplete work in production.
**When to use.** Default for most teams. Especially when CI pipeline duration is short (under 30 minutes) and observability is good enough to detect production issues quickly.
**Common problems.**
- Teams with infrequent releases find it hard to "hold" features for a release window. Feature flags are the answer, not branches.
- Teams without good test coverage shouldn't trunk-base — every commit hits production-shaped pipelines, so flaky tests block everyone.
**Worktree mapping.** Create: `aidlc engine worktree create --slug <bolt-slug> --base main`. Merge: `--target main --strategy squash`. Each Bolt = one squash commit on `main`.
**Parallel Bolts.** Cleanest fit. Multiple Bolts can be in flight simultaneously; each branches from current `main`, each merges back without rebase contention because squash flattens history at merge time.
### Execution runbook
When dispatched for trunk-based:
1. Read `## Way of Working` from the active space's `project.md`, `team.md`, then `org.md` per `shared/rules-reading.md`; use hardcoded defaults only if all three are empty.
2. Resolve flags: `--base main --target main --strategy squash` for the default; deviate only if `team.md` explicitly says otherwise.
3. **Create**: invoke `aidlc engine worktree create --slug <bolt-slug> --base main`.
4. **Merge** (after Bolt gate approval): caller must be on `main` at the main checkout. Invoke `aidlc engine worktree merge --slug <bolt-slug> --target main --strategy squash --message "<commit message>"`.
5. Return the JSON envelope per § Response contract back to the orchestrator.
### Failure modes
- **Dirty tree on merge.** Local uncommitted changes on `main`; tool errors with the git message verbatim. Orchestrator's halt-and-ask offers retry/abort. Worktree preserved on retry; user explicitly discards on abort.
- **Conflict on squash.** Squash conflicts with concurrent `main` motion (e.g. another Bolt landed first). Tool exits non-zero with `{status: "conflict", conflict_files, detail}`. Orchestrator quotes `detail` to the user.
- **Branch already exists.** Pre-audit error; the tool refuses to clobber. Orchestrator should run `discard` first (rare) or pick a different slug.
---
## GitHub Flow
```
main ────────────────────────────────────────►
▲ ▲ ▲ ▲
│ │ │ │ feature branches with PRs
│ │ │ │ (no time limit; can live longer than 1-2 days)
feat-A feat-B feat-C feat-D
```
**Shape.** `main` + indefinite feature branches. PRs merge to `main`. Branches can live longer than trunk-based — days to weeks for larger features. No `develop` or `release` branches.
**When to use.** Teams that want trunk-based discipline but need longer-lived feature branches. Open-source projects often use this.
**Common problems.**
- Branches that live too long accumulate merge debt. Discipline required to either land or close.
- Without feature flags, in-flight features block release of unrelated work.
**Worktree mapping.** Same base/target as trunk-based: `--base main --target main`. Strategy is usually `squash`, but teams that prefer to preserve the branch in history use `merge`. Team picks at affirmation; agent reads `team.md`.
### Execution runbook
When dispatched for GitHub Flow:
1. Read `## Way of Working` from the active space's `project.md`, `team.md`, then `org.md` (including any merge-style statement). The merge-strategy choice (squash vs merge) is what differs from trunk-based.
2. Resolve flags: `--base main --target main --strategy <squash|merge>` per affirmation; default to `squash`.
3. **Create**: `aidlc engine worktree create --slug <bolt-slug> --base main`.
4. **Merge**: `aidlc engine worktree merge --slug <bolt-slug> --target main --strategy <squash|merge> [--message "<msg>"]`. With `--strategy merge`, a no-fast-forward merge commit preserves the bolt branch's individual commits.
5. Return per § Response contract.
### Failure modes
- **Same as trunk-based**, plus:
- **Stale base for `--strategy merge`.** Long-lived bolt branches against a moving `main` produce conflicts. The tool reports the conflict envelope; the user resolves in the worktree (preserved on conflict) and re-invokes merge.
---
## GitFlow
```
main ────────────────────────────────►
▲ ▲
│ │ release/v1.0 release/v1.1
develop ─────┴───────────┴────────────────────►
▲ ▲ ▲ ▲ ▲
│ │ │ │ │ feature branches off develop
│ │ │ │ │
feat-A feat-B feat-C feat-D
▲
│ hotfix/v1.0.1 (off main, merged to both)
```
**Shape.** Two long-lived branches (`main` = production, `develop` = integration), plus `feature/*`, `release/*`, `hotfix/*` short-lived branches. Releases cut from `develop` → `release/*` → `main` with version tags.
**When to use.** Teams with strict release management — quarterly releases, regulated deployments, stable production while integration continues. Common in enterprise + financial services.
**Common problems.**
- Long-lived `develop` accumulates merge debt against `main` over a release cycle. Painful merges at release-cut time.
- Hotfixes require dual-merging (to both `main` and `develop`) — easy to miss the second merge.
**Worktree mapping.** Feature Bolts: `--base develop --target develop`. Hotfix Bolts: `--base main --target main` (and the operator merges to `develop` separately — out of scope for `aidlc-worktree`). Strategy is usually `merge` to preserve branch history; `squash` is also valid.
### Execution runbook
When dispatched for GitFlow:
1. Read the active space's `## Way of Working`. Look for the integration-branch name (`develop` is the convention; teams sometimes use `integration` or `next`).
2. For feature Bolts: `--base <integration> --target <integration> --strategy <merge|squash>`. Default to `merge`.
3. For hotfix Bolts (rare in Construction; usually triggered by an out-of-band stage): `--base main --target main --strategy merge`. The operator separately merges the hotfix back to `<integration>` after `aidlc-worktree merge` succeeds. Out of scope for the tool.
4. **Create**: `aidlc engine worktree create --slug <bolt-slug> --base <integration>`.
5. **Merge**: caller must be on `<integration>` at the main checkout. `aidlc engine worktree merge --slug <bolt-slug> --target <integration> --strategy <merge|squash>`.
6. Return per § Response contract; if hotfix, include `notes: "manual merge to <integration> required"` so the orchestrator surfaces the follow-up.
### Failure modes
- **`<integration>` branch missing locally.** Pre-audit error; tool refuses to invent the branch.
- **Wrong cwd on merge.** Defensive HEAD check fails: `expected branch <integration>, found <actual>`. Caller must `cd` to the main checkout and `git checkout <integration>` first.
- **Hotfix merge to second target forgotten.** Out-of-scope for `aidlc-worktree`; orchestrator's aidlc-pipeline-deploy-agent dispatch should always include the second-target reminder in `notes`.
---
## Release Branches
```
main ────────────────────────────────►
▲ ▲
│ │
release/v1.0 ┴──── (frozen for stabilisation) ──────►
▲ ▲
│ │ bug fixes only on release branch
│ │
fix-A fix-B
│
└──► merge release/v1.0 → main + tag v1.0.0
```
**Shape.** Trunk-based or GitHub Flow on `main`, with a release branch cut at code-freeze. Stabilisation work (bug fixes only) happens on the release branch; new features continue on `main`.
**When to use.** Teams shipping versioned software where release stability matters more than continuous deployment — desktop apps, embedded software, enterprise products with hard release dates.
**Common problems.**
- Bug fixes on release branch must be cherry-picked or merged back to `main` so they don't regress in the next release.
- Long stabilisation periods can block feature work waiting for the release branch to merge back.
**Worktree mapping.** Bolts on `main` use `--base main --target main`. Release-branch fix Bolts use `--base release/vX.Y --target release/vX.Y`. Strategy is `merge` typically (preserves the fix branch in history for traceability).
### Execution runbook
When dispatched for Release Branches:
1. Read the active space's `## Way of Working`. Look for the release-branch pattern (`release/vX.Y` is the convention).
2. Determine which line the Bolt belongs to from the Bolt's metadata (the orchestrator passes a `target_line: main | release/vX.Y` hint). Default to `main` when ambiguous.
3. **Create**: `aidlc engine worktree create --slug <bolt-slug> --base <line>`.
4. **Merge**: caller on `<line>` at the main checkout. `aidlc engine worktree merge --slug <bolt-slug> --target <line> --strategy merge`.
5. If the Bolt was a release-branch fix, include `notes: "consider cherry-pick to main"` in the response — the operator handles the cross-merge.
6. Return per § Response contract.
### Failure modes
- **Release branch missing locally.** Pre-audit error.
- **Bolt targeted release branch but main has diverged.** Bolt completes; orchestrator surfaces the cherry-pick reminder via the `notes` field.
- **Same as GitFlow** for wrong-cwd / dirty-tree / conflict cases.
---
## Monorepo
```
main ────────────────────────────────────────►
▲ ▲ ▲ ▲
│ │ │ │ feature branches with path-based scope
│ │ │ │
pkg-a/feat-1 pkg-b/feat-2 pkg-c/refactor shared/lib-update
```
**Shape.** Single repo holding multiple packages/services. Branches scoped by path (changes within `packages/auth/` are one Bolt; changes spanning packages need explicit cross-package coordination). Can run trunk-based, GitHub Flow, or any of the above on top.
**When to use.** Teams with multiple closely-coupled services that benefit from atomic cross-service changes. Tooling support required: Nx, Turborepo, Pants.
**Common problems.**
- CI must be path-aware (only test packages with changes). Monolithic CI defeats the purpose.
- Cross-package changes can't be parallelised cleanly — they require coordinated merges.
**Worktree mapping.** Same as the underlying strategy (trunk-based default). The path-awareness lives at the CI/test layer, not the worktree layer. Strategy is usually `squash` per package change.
### Execution runbook
When dispatched for Monorepo:
1. Resolve the underlying strategy (trunk-based default) per its runbook above.
2. The Bolt slug should encode the package scope (e.g. `auth-token-rotation` rather than `feature-1`) so `git worktree list` output stays diagnosable.
3. Create + merge identical to the underlying strategy.
4. Return per § Response contract.
### Failure modes
- **Cross-package Bolts.** When a Bolt's units span two packages, the merge succeeds but the cherry-pick / coordinate-with-other-package reminder is the operator's job. Surface in `notes` if known.
- **Same as the underlying strategy.**
---
## Response contract
When the orchestrator dispatches aidlc-pipeline-deploy-agent for a worktree create or merge, the agent invokes `aidlc-worktree` directly and reports the JSON envelope below back to the orchestrator. SKILL.md Step 0.5 / Step 6.75 then call `aidlc-worktree verify` as a deterministic backstop confirming the audit event landed.
### Create response (success)
```json
{
"emitted": "WORKTREE_CREATED",
"slug": "<bolt-slug>",
"worktree_path": "/abs/path/.aidlc/worktrees/bolt-<slug>",
"branch": "bolt-<slug>",
"base": "<base-branch>",
"audit_timestamp": "2026-05-18T12:34:56Z",
"notes": "<optional follow-up reminders for the orchestrator>"
}
```
### Merge response (success)
```json
{
"emitted": "WORKTREE_MERGED",
"slug": "<bolt-slug>",
"worktree_path": "/abs/path/.aidlc/worktrees/bolt-<slug>",
"target": "<target-branch>",
"strategy": "squash",
"commit_sha": "<sha>",
"audit_timestamp": "2026-05-18T12:34:56Z",
"notes": "<optional follow-up reminders>"
}
```
If a merge error carries `[merge-succeeded:<sha>]` and says the
`SWARM_SOURCE_MERGED` post-result audit row failed, the Git merge already
landed but no aggregate source authority exists. Preserve the worktree and do
not retry the same merge command. Restart the stage attempt, or use
`AIDLC_SKIP_SOURCE_FRESHNESS=1` only after explicit human approval.
### Merge response (conflict)
```json
{
"status": "conflict",
"slug": "<bolt-slug>",
"worktree_path": "/abs/path/.aidlc/worktrees/bolt-<slug>",
"conflict_files": ["src/foo.ts", "src/bar.ts"],
"detail": "Merge produced conflicts in worktree at <path>. Worktree preserved for inspection."
}
```
The orchestrator's halt-and-ask quotes the `detail` field verbatim. See `aidlc-common/protocols/stage-protocol-construction.md` § "Halt-and-ask on failure" and `skills/aidlc/SKILL.md` § "Halt-and-ask failure handling" for the full prompt shape and preservation invariant.
### Discard response
```json
{
"emitted": "WORKTREE_DISCARDED",
"slug": "<bolt-slug>",
"worktree_path": "/abs/path/.aidlc/worktrees/bolt-<slug>",
"reason": "agent-discard",
"audit_timestamp": "2026-05-18T12:34:56Z"
}
```
If the worktree was already gone (idempotent path), `emitted` is `null` and `reason` is `already-discarded` — no audit event is re-emitted.
---
## How AIDLC reads strategy from team practices
The dispatch protocol described in this section is implemented by **SKILL.md Step 0** (worktree create) and **Step 6.5** (worktree merge). `aidlc-bolt complete --merge` orchestrates around the dispatch (forkState merge-back, forkAudit merge-back) but does not call `aidlc-worktree merge` directly — the dispatch lives in SKILL.md prose.
When a Bolt starts (Step 0) or completes (Step 6.5), the orchestrator dispatches a Task call to **aidlc-pipeline-deploy-agent** with two inputs:
1. The resolved `## Way of Working` statement from `aidlc/spaces/<active-space>/memory/{project,team,org}.md` (fallback chain in `shared/rules-reading.md`).
2. The Bolt's metadata (slug, source branch, optional target-line hint for release-branch teams).
The agent reads this file (`branching-strategies.md`) as the menu, matches the team's stated strategy to one of the five above, picks the right `aidlc-worktree` flags, invokes the tool, and returns the response envelope per § Response contract.
If the team's stated strategy doesn't map cleanly to the menu (e.g. "we use a hybrid"), the agent picks the closest fit and notes the deviation in the response's `notes` field; the orchestrator surfaces it in the audit log.
If none of `project.md`, `team.md`, or `org.md` provides a branching practice, the agent applies hardcoded defaults — trunk-based with squash, base `main`, target `main` — and emits `PRACTICES_SECTION_EMPTY` (advisory-only).
---
## Quick decision matrix
| You want... | Use |
|---|---|
| Default for a new project | Trunk-Based |
| OSS-style PRs with longer-lived branches | GitHub Flow |
| Enterprise release management | GitFlow |
| Versioned releases with stabilisation periods | Release Branches |
| Multiple services in one repo | Monorepo (on top of one of the above) |
If unsure, choose Trunk-Based. It's the lowest-overhead strategy with the strongest CI/CD ecosystem support, and AIDLC's Construction worktrees are designed for it as the default.
@@ -0,0 +1,91 @@
# CI/CD Pipeline Patterns
Patterns for building reliable continuous integration and continuous delivery pipelines.
## CI Pipeline Stages
A well-structured CI pipeline runs these stages in order:
1. **Lint** — Code formatting and style checks (ESLint, Prettier, Ruff, Black). Catch trivial issues before deeper analysis. Fast (< 30 seconds).
2. **Build** — Compile, transpile, or bundle the application. Verify the code produces valid artifacts. For CDK: `cdk synth`.
3. **Unit Test** — Run fast, isolated tests. Target: < 3 minutes. Fail the pipeline on any test failure.
4. **Static Analysis / Security Scan** — SAST, IaC scanning, dependency audit. Report findings; gate on severity thresholds.
5. **Integration Test** — Test against real dependencies (databases, queues) using testcontainers or LocalStack. Target: < 10 minutes.
6. **Package** — Build deployable artifacts: Docker images, Lambda ZIPs, CloudFormation templates. Tag with commit SHA and semantic version.
## CD Pipeline Patterns
### Continuous Delivery
Every commit that passes CI is deployable, but a human triggers the production deployment. Use manual approval gates in CodePipeline or GitHub Actions environments.
### Continuous Deployment
Every commit that passes CI and staging validation deploys to production automatically. Requires high confidence in test coverage and automated rollback.
**Recommendation**: Start with continuous delivery. Graduate to continuous deployment as test maturity and observability improve.
## Branch Strategies
> See `branching-strategies.md` in this directory for the deeper menu aidlc-pipeline-deploy-agent surveys at Bolt-merge dispatch — including diagrams, common problems, AIDLC worktree mapping per strategy, and parallel-Bolts notes for each. The summaries below cover the CI-flow context.
### Trunk-Based Development (Preferred)
- All developers commit to `main` (or short-lived feature branches merged within 1-2 days).
- Feature flags gate incomplete work. No long-lived branches.
- CI runs on every push to main. CD deploys from main.
- Benefits: Fewer merge conflicts, faster feedback, simpler pipeline.
### GitFlow
- `main` (production), `develop` (integration), `feature/*`, `release/*`, `hotfix/*`.
- Suitable for teams with infrequent releases or strict release management.
- Drawback: Long-lived branches cause painful merges and delayed integration.
### GitHub Flow
- `main` + short-lived feature branches. PRs merge to main.
- Simpler than GitFlow but still relies on branching. Good middle ground.
## Quality Gates
Define explicit pass/fail criteria that block pipeline progression:
| Gate | Criteria | Stage |
|------|----------|-------|
| Lint | Zero lint errors | Pre-build |
| Test coverage | >= 80% line coverage, no decrease | Post-test |
| Security scan | Zero Critical/High findings | Post-scan |
| Integration tests | 100% pass rate | Post-integration |
| Manual approval | Tech lead or product owner sign-off | Pre-production |
| Smoke tests | Critical path tests pass in production | Post-deploy |
## Artifact Management
- **Amazon ECR**: Store Docker images. Use immutable tags (commit SHA), not `latest`.
- **AWS CodeArtifact**: Host npm, pip, Maven packages. Proxy upstream registries for caching and security.
- **S3**: Store Lambda deployment ZIPs, CloudFormation templates, and build outputs.
- Tag every artifact with: commit SHA, build number, branch, timestamp.
- Set lifecycle policies: keep the last 30 tagged images; expire untagged images after 7 days.
## Pipeline as Code
- Define pipelines in version-controlled files, not UI configurations.
- **GitHub Actions**: `.github/workflows/*.yml`. Matrix builds for multi-version testing.
- **AWS CodePipeline + CodeBuild**: `buildspec.yml` for build steps; pipeline defined in CDK or CloudFormation.
- **Reusable workflows**: Extract common steps (lint, test, deploy) into shared workflow templates. Avoid copy-paste across repos.
## Monorepo vs Polyrepo CI Strategies
### Monorepo
- Use path-based triggers: only build/test services whose files changed.
- Tools: Nx, Turborepo, Pants. These understand dependency graphs and skip unchanged packages.
- Shared CI steps (lint, security scan) run once; service-specific steps run conditionally.
### Polyrepo
- Each repo has its own pipeline. Simpler per-repo but harder to coordinate cross-service changes.
- Use contract testing (Pact) to verify compatibility across repos without a monolithic integration test.
- Automate dependency updates across repos with Renovate or a custom bot.
## Pipeline Performance
- Target: commit to production in under 30 minutes (excluding manual approval).
- Parallelize independent stages (lint + unit test, SAST + dependency scan).
- Cache dependencies aggressively (npm cache, pip cache, Docker layer caching).
- Use spot instances or reserved capacity for build workers to reduce cost.
- Monitor pipeline duration as a team metric; alert when it exceeds thresholds.
@@ -0,0 +1,100 @@
# Deployment Strategies
Patterns for releasing software safely with minimal risk to users and the ability to roll back quickly.
## Blue/Green Deployment
**How it works**: Maintain two identical environments (blue = current, green = new). Deploy the new version to green. Switch traffic from blue to green at the load balancer or DNS level. Keep blue running as an instant rollback target.
**AWS Implementation**:
- ECS: Use CodeDeploy with `ECS` deployment type. Two target groups on an ALB; CodeDeploy shifts traffic.
- Lambda: Use aliases with weighted traffic shifting (`AWS::Lambda::Alias` with `RoutingConfig`).
- Elastic Beanstalk: Swap environment URLs.
**Advantages**: Instant rollback (repoint to blue), full environment validation before switch.
**Drawbacks**: Double infrastructure cost during deployment window. Database schema must be backward-compatible.
## Canary Releases
**How it works**: Route a small percentage of traffic (1-5%) to the new version. Monitor error rates, latency, and business metrics. Gradually increase traffic if healthy; roll back if anomalies are detected.
**AWS Implementation**:
- CodeDeploy with Lambda or ECS: Built-in canary configurations (`Canary10Percent5Minutes`, `Linear10PercentEvery1Minute`).
- API Gateway: Canary release on stage with percentage-based traffic split.
- CloudWatch alarms trigger automatic rollback on metric breaches.
**Advantages**: Limits blast radius. Detects issues with real traffic before full rollout.
**Drawbacks**: Requires robust monitoring. Stateful services need careful handling.
## Rolling Updates
**How it works**: Replace instances/tasks in batches. New version replaces a subset while the rest continue serving. Repeat until all instances run the new version.
**AWS Implementation**:
- ECS: Default deployment strategy. Configure `minimumHealthyPercent` and `maximumPercent`.
- EC2 Auto Scaling: Rolling update policy with `MinInstancesInService`.
**Advantages**: No extra infrastructure cost. Gradual rollout.
**Drawbacks**: Mixed versions during deployment (ensure backward compatibility). Slower rollback (redeploy the previous version).
## A/B Testing
**How it works**: Route specific user segments to different versions based on attributes (user ID, region, account type). Measure business outcomes (conversion, engagement) to decide which version wins.
**Distinction from canary**: A/B testing serves different experiences intentionally for experimentation; canary is a deployment safety mechanism.
**AWS Implementation**: CloudWatch Evidently for feature experiments with statistical analysis. CloudFront + Lambda@Edge for routing based on cookies or headers.
## Feature Flags
**How it works**: Deploy code with features wrapped in conditional flags. Toggle features on/off without redeployment.
**AWS Implementation**:
- **AppConfig**: Feature flags with validation, gradual rollout, and automatic rollback.
- **CloudWatch Evidently**: Feature flags with built-in A/B testing and metric tracking.
**Best Practices**:
- Use feature flags for incomplete features merged to main (trunk-based development enabler).
- Clean up flags after full rollout; stale flags become technical debt.
- Categorize flags: release flags (temporary), ops flags (kill switches), experiment flags (A/B tests).
- Never put secrets or sensitive config in feature flags.
## Rollback Strategies
### Automated Rollback
- Configure CloudWatch alarms on error rate, latency p99, and 5xx count.
- CodeDeploy automatically rolls back when alarms trigger during deployment.
- Lambda: Revert alias to the previous version instantly.
- ECS: CodeDeploy reroutes traffic back to the original target group.
### Manual Rollback
- Keep the previous artifact (Docker image, Lambda ZIP) tagged and deployable.
- Document the rollback procedure as a runbook: which commands, in what order, who approves.
- Practice rollbacks regularly; an untested rollback is not a rollback plan.
## Database Migration During Deployment
The hardest part of zero-downtime deployment is schema changes. Follow the **expand-contract** pattern:
1. **Expand**: Add new columns/tables/indexes. Do not remove or rename existing ones. Deploy application code that writes to both old and new schemas.
2. **Migrate**: Backfill data from old schema to new schema.
3. **Contract**: After all application versions use the new schema, remove old columns/tables in a later release.
Never run destructive schema changes (DROP COLUMN, rename) in the same deployment as the application change.
## Zero-Downtime Deployment Checklist
- [ ] Load balancer health checks configured with appropriate thresholds
- [ ] Connection draining enabled (deregistration delay: 30-120 seconds)
- [ ] Graceful shutdown handling in application (finish in-flight requests, close DB connections)
- [ ] Database schema changes are backward-compatible
- [ ] Rollback plan tested and documented
- [ ] Monitoring and alarms in place before deployment starts
- [ ] Pre-deployment smoke tests pass in staging
## Deployment Windows and Freeze Periods
- Define standard deployment windows (e.g., Tuesday-Thursday 10:00-16:00 local time) when the team is available to monitor.
- Enforce code freeze periods during high-traffic events (Black Friday, product launches) or compliance windows (end of fiscal quarter).
- Emergency hotfixes bypass freeze periods but require incident commander approval and post-deployment review.
- Automate freeze enforcement in the pipeline (reject deployments outside approved windows unless override flag is set).
@@ -0,0 +1,127 @@
# Functional Design Guide
## Business Logic Modeling
### Logic Decomposition Approach
Break complex business logic into composable layers:
1. **Input Validation Layer**: Verify data format, ranges, required fields
2. **Business Rule Layer**: Apply domain rules, constraints, and calculations
3. **State Transition Layer**: Manage entity lifecycle and valid state changes
4. **Side Effect Layer**: Trigger notifications, audit logs, integrations
### Business Rule Specification Format
For each rule, document (the `BR{group}.{seq}` ID format the traceability sensor recognizes, e.g. `BR1.1`):
```
Rule ID: BRx.y
Name: [descriptive name]
Trigger: [when this rule is evaluated]
Condition: [the logical expression]
Action: [what happens when condition is true]
Exception: [what happens when condition is false or invalid]
Priority: [execution order when multiple rules apply]
Source: [requirement ID or stakeholder that defined this rule]
```
### Rule Conflict Resolution
When multiple rules apply to the same operation:
- **Priority ordering**: Higher-priority rules execute first
- **First-match wins**: Stop evaluating after the first matching rule
- **All-match accumulate**: Apply all matching rules (must be non-contradictory)
- Always document the chosen strategy per rule group
## Domain Entity Design
### Entity Identification Checklist
An entity should be modeled when:
- It has a unique identity that persists over time
- It has a lifecycle with distinct states
- Multiple parts of the system reference it
- It has business rules governing its behavior
- It participates in relationships with other entities
### Entity Specification Template
| Attribute | Type | Required | Constraints | Default | Notes |
|-----------|------|----------|-------------|---------|-------|
| id | UUID | Yes | System-generated | Auto | Primary identifier |
| status | Enum | Yes | [valid values] | Initial | See state machine |
| ... | ... | ... | ... | ... | ... |
### Relationship Types and Design Rules
- **One-to-One**: Embed or separate based on access patterns (always queried together = embed)
- **One-to-Many**: Parent owns the collection; child references parent by ID
- **Many-to-Many**: Use a junction entity with its own attributes (timestamps, status)
- Always define cascade behavior: what happens to children when parent is deleted?
- Document referential integrity constraints explicitly
## Business Rule Specification Patterns
### Calculation Rules
For computed values, specify:
- **Formula**: The calculation expression with variable definitions
- **Precision**: Rounding rules, decimal places, currency handling
- **Edge cases**: Division by zero, overflow, null inputs
- **Examples**: At least 3 worked examples with inputs and expected output
### Validation Rules
For each input field:
| Field | Type | Required | Min | Max | Pattern | Custom Rule |
|-------|------|----------|-----|-----|---------|-------------|
| email | string | Yes | 5 | 254 | RFC 5322 | Must be unique |
| amount | decimal | Yes | 0.01 | 999999.99 | 2 decimal places | Must not exceed account balance |
### Authorization Rules
Document access control per operation:
```
Operation: [Create/Read/Update/Delete] [Entity]
Allowed Roles: [role list]
Additional Conditions: [ownership check, status check, time window]
Denied Response: [error code and message]
Audit: [whether to log access attempts]
```
## Workflow Design Methodology
### Workflow Specification Template
For each business workflow:
1. **Trigger**: What initiates the workflow (user action, scheduled event, external signal)
2. **Preconditions**: What must be true before the workflow can start
3. **Steps**: Ordered list of actions with decision points
4. **Actors**: Who or what performs each step (user, system, external service)
5. **Postconditions**: What must be true when the workflow completes successfully
6. **Error Paths**: What happens when each step fails
7. **Timeout Behavior**: What happens if the workflow stalls at any step
### State Machine Design
For entities with complex lifecycles:
| Current State | Event | Guard Condition | Next State | Actions |
|--------------|-------|-----------------|------------|---------|
| Draft | Submit | All required fields populated | Pending Review | Notify reviewer, log event |
| Pending Review | Approve | Reviewer has authority | Approved | Notify submitter, update timestamps |
| Pending Review | Reject | Rejection reason provided | Draft | Notify submitter with reason |
| Approved | Activate | Start date reached | Active | Enable functionality |
| Active | Expire | End date reached | Expired | Disable functionality, notify owner |
| Any | Cancel | Cancellation policy met | Cancelled | Notify stakeholders, release resources |
### State Machine Validation Rules
- Every state must be reachable from the initial state
- Every non-terminal state must have at least one outgoing transition
- Terminal states (Cancelled, Expired, Completed) have no outgoing transitions
- No implicit transitions -- every state change requires an explicit event
- Guard conditions must be testable (no subjective criteria)
## Functional Design Document Structure
Organize each functional design specification as:
1. **Overview**: Purpose and scope of the function
2. **Entities**: Data model with attributes and relationships
3. **Business Rules**: Complete rule set with priorities
4. **Workflows**: Step-by-step processes with decision points
5. **State Machines**: Lifecycle diagrams for stateful entities
6. **Interfaces**: Input/output specifications for each operation
7. **Validation**: Field-level and cross-field validation rules
8. **Error Catalog**: All error conditions with codes and messages
9. **Traceability**: Mapping back to requirements (FR-NNN references)
@@ -0,0 +1,105 @@
# Market Research Methods
## Purpose
Structured approaches for understanding the competitive landscape, sizing the opportunity, and making informed build-vs-buy decisions before committing engineering effort.
## Competitive Analysis Framework
### Step 1: Identify Competitors
- **Direct competitors**: Same product, same market (e.g., Slack vs Teams)
- **Indirect competitors**: Different product, same problem (e.g., Slack vs email)
- **Potential competitors**: Adjacent players who could enter your space
### Step 2: Feature Comparison Matrix
| Capability | Our Product | Competitor A | Competitor B | Competitor C |
|------------|-------------|-------------|-------------|-------------|
| Core Feature 1 | | | | |
| Core Feature 2 | | | | |
| Pricing Model | | | | |
| Target Segment | | | | |
| Integration Ecosystem | | | | |
| Key Differentiator | | | | |
Rate each: Strong / Adequate / Weak / Absent
### Step 3: Positioning Map
Plot competitors on two axes representing the dimensions that matter most to your target users (e.g., Ease of Use vs Power, Price vs Completeness). Identify underserved quadrants.
## SWOT Analysis Template
| | Helpful | Harmful |
|---|---------|---------|
| **Internal** | **Strengths**: What do we do well? What unique resources do we have? | **Weaknesses**: Where do we lack capability? What do competitors do better? |
| **External** | **Opportunities**: What market trends favor us? What unmet needs exist? | **Threats**: What could disrupt us? What are competitors planning? |
### Making SWOT Actionable
- **Strengths + Opportunities** = Strategies to pursue aggressively
- **Weaknesses + Threats** = Risks to mitigate or defend against
- **Strengths + Threats** = Defensive strategies leveraging advantages
- **Weaknesses + Opportunities** = Investments needed to capitalize
## Porter's Five Forces
Assess industry attractiveness by analyzing:
1. **Threat of New Entrants**: How easy is it to enter this market? (Low barriers = high threat)
2. **Bargaining Power of Suppliers**: How dependent are you on key suppliers/platforms? (Single cloud provider = high power)
3. **Bargaining Power of Buyers**: Can customers easily switch? (Low switching costs = high power)
4. **Threat of Substitutes**: Can the problem be solved differently? (Manual process, different tech)
5. **Competitive Rivalry**: How intense is existing competition? (Many similar products = high rivalry)
Rate each force High/Medium/Low. High forces compress margins and reduce attractiveness.
## Market Sizing: TAM / SAM / SOM
### Definitions
- **TAM (Total Addressable Market)**: Total revenue if you captured 100% of the market globally
- **SAM (Serviceable Addressable Market)**: Portion of TAM your product/model can realistically serve (geography, segment, channel)
- **SOM (Serviceable Obtainable Market)**: Realistic market share you can capture in 2-3 years given competition and resources
### Top-Down Calculation
Start with industry reports, narrow by filters:
```
TAM = Total industry revenue for category
SAM = TAM x % in your geography x % in your segment
SOM = SAM x realistic capture rate (typically 1-5% for new entrants)
```
### Bottom-Up Calculation (More Reliable)
```
SOM = Target customers x Average deal size x Purchase frequency
```
Use both methods and compare. If they diverge wildly, investigate assumptions.
## Build vs Buy Assessment
### Evaluation Criteria
| Factor | Build | Buy | Score (Build -2 to +2) |
|--------|-------|-----|----------------------|
| Core differentiator? | Build if it is your competitive advantage | Buy commodity capabilities | |
| Time to market | Months to years | Days to weeks | |
| Total cost (3-year) | Dev + maintenance + opportunity cost | License + integration + vendor risk | |
| Customization needed | High customization favors build | Standard needs favor buy | |
| Data sensitivity | Full control | Vendor data handling policies | |
| Team expertise | Requires in-house skills | Transfers complexity to vendor | |
| Maintenance burden | Ongoing responsibility | Vendor handles updates | |
### Decision Rule
If the capability is not your core differentiator and a mature vendor solution exists, default to Buy. Build only when the capability is central to your competitive advantage or when vendor solutions fundamentally cannot meet your requirements.
## Trend Analysis Methods
### Approaches
- **Technology radar**: Categorize emerging technologies as Adopt / Trial / Assess / Hold
- **Hype cycle mapping**: Identify where relevant technologies sit on the Gartner Hype Cycle
- **Customer signal analysis**: Mine support tickets, feature requests, and churn reasons for emerging patterns
- **Adjacent market monitoring**: Track what is happening in related industries that may spill over
### Validation
Trends are hypotheses. Validate with:
- Customer interviews confirming the trend affects their buying decisions
- Revenue data showing market movement (not just media coverage)
- Competitor investment signals (hiring patterns, acquisitions, product launches)
@@ -0,0 +1,104 @@
# Prioritization Frameworks
## Purpose
Structured methods for deciding what to build first. Every framework encodes trade-offs — choose the one that matches your decision context.
## Framework Selection Guide
| Framework | Best For | Complexity | Stakeholder Buy-in |
|-----------|----------|------------|---------------------|
| MoSCoW | Fixed-scope releases, contract work | Low | Easy |
| WSJF | Lean/SAFe teams, flow-based delivery | Medium | Medium |
| RICE | Data-driven product teams | Medium | High |
| Kano Model | UX-focused products, feature differentiation | High | High |
## MoSCoW Method
### Categories
- **Must Have** — Non-negotiable for this release. System is unusable without it.
- **Should Have** — Important but not critical. Painful to omit but workarounds exist.
- **Could Have** — Desirable. Include if time/budget allows.
- **Won't Have (this time)** — Explicitly out of scope. Prevents scope creep.
### Rules of Thumb
- Must Haves should not exceed 60% of planned capacity
- If everything is a Must Have, you have not prioritized — push back
- "Won't Have" is the most valuable category: it creates clarity
## WSJF (Weighted Shortest Job First)
### Formula
```
WSJF = Cost of Delay / Job Duration
Cost of Delay = User-Business Value + Time Criticality + Risk Reduction
```
### Scoring
Rate each factor 1-10 relative to other items in the backlog (Fibonacci scale: 1, 2, 3, 5, 8, 13 also works). Divide by estimated duration (relative size).
### When WSJF Wins
- When you need to maximize value throughput, not just value
- Small high-value items naturally float to the top
- Forces teams to consider the cost of NOT doing something now
## RICE Scoring
### Formula
```
RICE Score = (Reach x Impact x Confidence) / Effort
```
### Component Definitions
- **Reach**: How many users/transactions affected per quarter (use real numbers)
- **Impact**: Per-user effect (3 = massive, 2 = high, 1 = medium, 0.5 = low, 0.25 = minimal)
- **Confidence**: Data quality (100% = high confidence with data, 80% = medium, 50% = low/gut feel)
- **Effort**: Person-months of work (use whole numbers)
### Example
| Feature | Reach | Impact | Confidence | Effort | RICE |
|---------|-------|--------|------------|--------|------|
| Search autocomplete | 10,000 | 2 | 80% | 2 | 8,000 |
| Admin dashboard | 50 | 3 | 100% | 4 | 37.5 |
| Onboarding wizard | 5,000 | 3 | 50% | 3 | 2,500 |
### Key Insight
Confidence is the honesty check. Low-confidence, high-impact ideas should be validated (spike, prototype, user test) before committing full build effort.
## Kano Model
### Category Definitions
- **Basic (Must-Be)**: Expected features. Absence causes dissatisfaction; presence does not delight. (e.g., login works, pages load)
- **Performance (One-Dimensional)**: More is better. Linear relationship between investment and satisfaction. (e.g., speed, storage)
- **Excitement (Attractive)**: Unexpected features that delight. Absence does not disappoint. (e.g., smart suggestions, delightful animations)
- **Indifferent**: Users do not care either way. Stop investing here.
- **Reverse**: Features some users actively dislike. (e.g., forced tutorials, auto-play)
### Classification Method
Ask two questions per feature:
1. "How would you feel if this feature were present?" (functional)
2. "How would you feel if this feature were absent?" (dysfunctional)
Map answers (Like, Expect, Neutral, Tolerate, Dislike) to the Kano evaluation table.
## Handling Ties and Disputes
### When Scores Are Equal
1. Prefer the item with higher confidence — less risk
2. Prefer the item that unblocks other work — multiplier effect
3. Prefer the item with shorter time-to-value — faster feedback loop
### When Stakeholders Disagree
- Make the framework transparent — show the math, not just the result
- Separate "I want this" from "users need this" with data (analytics, support tickets, user research)
- Use time-boxing: "We will revisit in 2 sprints with usage data"
- Escalation path: Product owner makes the final call, documents rationale
## Decision Matrix Template
| Item | Business Value (1-5) | User Impact (1-5) | Effort (1-5, inverted) | Risk (1-5) | Weighted Score |
|------|---------------------|-------------------|----------------------|------------|----------------|
| Feature A | | | | | |
| Feature B | | | | | |
Assign weights to each column based on current strategic priorities (e.g., Business Value x2 if revenue-focused quarter).
@@ -0,0 +1,89 @@
# Product Guide
## User Story Format
Standard format: **As a [persona], I want [action], so that [benefit].**
### INVEST Criteria Checklist
Every story must pass all six criteria:
- **I - Independent**: Can be developed and delivered without depending on another story
- **N - Negotiable**: Details can be discussed; it is not a rigid contract
- **V - Valuable**: Delivers identifiable value to a user or stakeholder
- **E - Estimable**: Team can estimate effort (if not, the story needs decomposition or spike)
- **S - Small**: Completable within a single iteration (1-5 days of implementation work)
- **T - Testable**: Has concrete acceptance criteria that can be verified
### Story Decomposition Patterns
When a story is too large (epic), split using these strategies:
1. **By workflow step**: Login -> Browse -> Select -> Purchase -> Confirm
2. **By data variation**: Handle text input / Handle file upload / Handle image upload
3. **By business rule**: Basic validation / Advanced validation / Cross-field validation
4. **By interface**: Web UI / API endpoint / Admin panel
5. **By operation**: Create / Read / Update / Delete (but prefer vertical slices)
## Persona Development
For each persona, define:
```
Name: [descriptive name, e.g., "Alex the Admin"]
Role: [their role in the system]
Goals: [what they want to accomplish, 2-3 items]
Pain Points: [current frustrations, 2-3 items]
Tech Comfort: [low / medium / high]
Frequency: [how often they use the system]
```
Every story must reference a defined persona. If a story does not fit any persona, either the story is wrong or a persona is missing.
## Prioritization Frameworks
### MoSCoW (preferred for MVP definition)
- **Must Have**: System is unusable without this. If removed, the product fails its core purpose.
- **Should Have**: Important but not critical. Workarounds exist. Include if time permits.
- **Could Have**: Desirable. Enhances experience but not expected in first release.
- **Won't Have (this time)**: Explicitly out of scope. Documented for future consideration.
### RICE Scoring (preferred for backlog ranking)
Score = (Reach x Impact x Confidence) / Effort
- **Reach**: How many users/sessions affected per time period (use real numbers)
- **Impact**: How much it moves the needle (3=massive, 2=high, 1=medium, 0.5=low, 0.25=minimal)
- **Confidence**: How sure are you about estimates (100%/80%/50%)
- **Effort**: Person-days of work (all roles combined)
## MVP Definition Criteria
The MVP must:
1. Solve the core problem for the primary persona
2. Include all Must Have stories and no Could/Won't stories
3. Be deployable and usable without manual workarounds
4. Include basic error handling (not just happy paths)
5. Meet minimum NFRs for security and data integrity
6. Be demonstrable to stakeholders in under 10 minutes
## Workflow Planning Structure
Organize stories into iterations:
```
Iteration 0 (Foundation): Infrastructure, auth, core data model
Iteration 1 (Core Value): Primary user workflow end-to-end
Iteration 2 (Completeness): Secondary workflows, edge cases, admin features
Iteration 3 (Polish): Performance optimization, UX refinement, advanced features
```
For each iteration, define:
- Entry criteria (what must be complete before starting)
- Stories included (with dependency order)
- Exit criteria (what "done" looks like for this iteration)
- Demo scenario (how to showcase the increment)
## Story Mapping Layout
Arrange stories in a 2D map:
- **Horizontal axis**: User journey steps (left to right, in sequence)
- **Vertical axis**: Priority (top = must-have, bottom = nice-to-have)
- **Horizontal line**: MVP boundary (everything above the line is MVP)
This visualization makes it easy to spot:
- Missing journey steps (vertical gaps)
- Over-invested areas (too many stories in one column)
- Dependency chains (stories that must be above others)
@@ -0,0 +1,86 @@
# Requirements Elicitation Techniques
## Purpose
Systematic methods for discovering, capturing, and validating what stakeholders truly need — not just what they initially say they want.
## Technique Selection Guide
| Technique | Best For | Effort | Fidelity |
|-----------|----------|--------|----------|
| Stakeholder Interviews | Deep understanding of individual needs | Medium | High |
| Workshops | Consensus building, conflict resolution | High | High |
| Observation | Understanding actual vs stated workflows | Medium | Very High |
| Document Analysis | Existing system understanding, compliance | Low | Medium |
| Prototyping | Validating assumptions, UI-heavy features | High | Very High |
## Stakeholder Interviews
### Preparation
- Research the stakeholder's role, responsibilities, and known pain points
- Prepare 8-12 open-ended questions; plan for 45-60 minutes
- Share agenda in advance so they can prepare examples
### Interview Question Templates
- "Walk me through a typical day when you [process]. What frustrates you most?"
- "If you could change one thing about the current system, what would it be and why?"
- "When [process] goes wrong, what happens? Who gets impacted?"
- "How do you measure success in [area]? What metrics matter?"
- "Show me how you currently do [task] — what workarounds have you developed?"
- "Who else should I talk to about this?"
### Common Pitfalls
- **Leading questions**: "Don't you think X would be better?" forces agreement
- **Assumption bias**: Projecting your solution onto their problem
- **HiPPO effect**: Letting the Highest-Paid Person's Opinion dominate
- **Survivorship bias**: Only talking to current users, ignoring churned users
- **Premature solutioning**: Jumping to "we could build X" before understanding the problem
## Workshops
### When to Use
- Multiple stakeholders with conflicting priorities
- Cross-functional alignment needed (e.g., sales vs engineering vs support)
- Time-boxed discovery needed (compressed timeline)
### Workshop Format
1. **Context setting** (10 min) — Problem statement, goals, ground rules
2. **Individual ideation** (10 min) — Silent sticky-note brainstorming prevents groupthink
3. **Share and cluster** (15 min) — Group similar ideas, identify themes
4. **Dot voting** (5 min) — Each participant gets 3 votes to prioritize
5. **Deep dive** (30 min) — Discuss top-voted items, capture acceptance criteria
6. **Wrap-up** (5 min) — Summarize decisions, assign follow-ups
## Observation (Contextual Inquiry)
### Method
- Watch users perform real tasks in their actual environment
- Ask "why" when you see unexpected behavior — workarounds reveal unmet needs
- Note environmental factors: interruptions, tool switching, manual data entry
### Key Insight
Users often cannot articulate their workflow because it is habitual. Observation reveals the gap between "what they say they do" and "what they actually do."
## Document Analysis
### Sources to Review
- Existing system documentation, help desk tickets, bug reports
- Regulatory requirements, compliance standards, audit findings
- Competitor product documentation and reviews
- Current process flowcharts, SOPs, training materials
## Prototyping for Requirements
### Progression
1. **Paper sketches** — Validate concepts in minutes, discard freely
2. **Clickable wireframes** — Test navigation and flow logic
3. **Functional prototypes** — Validate complex interactions, data-dependent UIs
### Rule of Thumb
Prototype the riskiest assumption first. If users struggle with the core concept in a paper sketch, building a functional prototype wastes effort.
## Validation Checklist
- [ ] Each requirement traces to at least one stakeholder need
- [ ] Requirements are testable (clear pass/fail criteria exist)
- [ ] No orphan requirements (requirements with no user or business justification)
- [ ] Conflicts between stakeholders are explicitly resolved and documented
- [ ] Non-functional requirements (performance, security, scale) are captured alongside functional ones
@@ -0,0 +1,88 @@
# Requirements Guide
## Requirement Types
### Functional Requirements (FR)
Define what the system must do. Format: "The system shall [verb] [object] [condition]."
- User-facing behavior (inputs, outputs, interactions)
- Business rules and logic (calculations, validations, state transitions)
- Data requirements (entities, relationships, lifecycle)
- Integration requirements (external systems, APIs, data feeds)
### Non-Functional Requirements (NFR)
Define how the system must perform. Must be quantifiable.
- **Performance**: Response time < Xms for Y% of requests under Z concurrent users
- **Availability**: X% uptime measured over a rolling 30-day window
- **Scalability**: Support X to Y concurrent users with linear resource scaling
- **Security**: Authentication method, authorization model, data encryption at rest/in transit
- **Usability**: WCAG level, maximum clicks to complete core workflow
- **Maintainability**: Code coverage target, deployment frequency target
### Constraints
Non-negotiable boundaries imposed externally:
- Technology mandates (must use AWS, must use React, must support IE11)
- Regulatory compliance (GDPR, HIPAA, SOC2, PCI-DSS)
- Budget and timeline limitations
- Integration compatibility with existing systems
### Assumptions
Believed-true conditions that have not been validated:
- Always document assumptions explicitly
- Assign an owner responsible for validating each assumption
- Track assumption status (unvalidated, confirmed, invalidated)
## Elicitation Techniques
Use these in order of preference for AI-DLC:
1. **Document analysis** -- Read existing docs, READMEs, wikis, API specs, database schemas
2. **Structured questioning** -- Ask targeted questions using the completeness checklist below
3. **Scenario walkthrough** -- Walk through key user journeys step by step
4. **Constraint identification** -- Ask "What must NOT happen?" and "What are the limits?"
5. **Edge case probing** -- For each requirement, ask "What happens when [unusual condition]?"
## Acceptance Criteria Pattern
Use Given/When/Then (Gherkin) format:
```
Given [precondition or initial state]
When [action or trigger]
Then [expected outcome]
And [additional outcomes if needed]
```
Each requirement should have:
- At least 1 happy-path scenario
- At least 1 error/edge-case scenario
- Boundary values for any numeric constraints
## Completeness Analysis Checklist
For every system, verify coverage of:
- [ ] User authentication and authorization
- [ ] Data input validation and sanitization
- [ ] Error handling and user-facing error messages
- [ ] Data persistence and retrieval (CRUD for each entity)
- [ ] Search and filtering capabilities
- [ ] Pagination for list views
- [ ] Audit logging for sensitive operations
- [ ] Notification/alerting requirements
- [ ] Data export/import capabilities
- [ ] Concurrent access and conflict resolution
- [ ] Session management and timeout behavior
- [ ] Offline/degraded mode behavior (if applicable)
- [ ] Localization and internationalization (if applicable)
- [ ] Accessibility requirements
- [ ] Data retention and deletion policies
## Traceability Matrix Format
```
| Req ID | Description | Priority | Status | Design Ref | Unit Ref | Test Ref |
|--------|-------------|----------|--------|------------|----------|----------|
| FR-001 | [summary] | Must | Approved | AD-003 | U-007 | TC-012 |
```
Every cell should be filled. Empty cells indicate gaps:
- Empty Design Ref: requirement not yet designed
- Empty Unit Ref: requirement not yet assigned for implementation
- Empty Test Ref: requirement not yet covered by test plan
@@ -0,0 +1,123 @@
# User Story Patterns
## Purpose
Write user stories that are small enough to deliver quickly, valuable enough to justify the work, and clear enough to test without ambiguity.
## Standard Format
```
As a [role], I want to [action], so that [benefit].
```
The "so that" clause is the most important part — it forces articulation of value. If you cannot state the benefit, question whether the story is needed.
## INVEST Criteria
Every story should satisfy all six criteria:
- **Independent** — Can be developed and delivered without depending on other stories
- **Negotiable** — Details are open to discussion; the story is not a contract
- **Valuable** — Delivers value to the user or business (not "set up database")
- **Estimable** — Team can reasonably estimate the effort required
- **Small** — Completable within a single sprint (ideally 1-3 days of work)
- **Testable** — Clear acceptance criteria define "done" without ambiguity
## Story Splitting Patterns
When a story is too large, split it using one of these patterns:
### 1. Workflow Steps
Split a multi-step process into individual steps.
- "User completes checkout" becomes: add to cart, enter shipping, enter payment, confirm order, receive confirmation
### 2. Business Rule Variations
Isolate each business rule into its own story.
- "Calculate shipping cost" becomes: flat rate domestic, weight-based domestic, international shipping, free shipping threshold
### 3. Data Variations
Split by the type or complexity of data handled.
- "Import customer data" becomes: import from CSV, import from API, import with deduplication, import with validation errors
### 4. Interface Variations
Split by platform or interaction mode.
- "User views dashboard" becomes: desktop layout, mobile responsive layout, keyboard-accessible version
### 5. Operations / CRUD
Split create, read, update, and delete into separate stories.
- Start with Read (simplest, immediate value), then Create, then Update, then Delete
### 6. Performance / Scale
Separate "make it work" from "make it fast."
- Story 1: "User searches products (basic query, < 1000 results)"
- Story 2: "User searches products with sub-200ms response across 1M+ catalog"
## Acceptance Criteria: Given/When/Then
### Format
```
Given [precondition / initial state]
When [action / trigger]
Then [expected outcome / observable result]
```
### Example
```
Story: As a customer, I want to reset my password, so that I can regain access to my account.
AC 1:
Given I am on the login page
When I click "Forgot password" and enter my registered email
Then I receive a password reset link within 5 minutes
AC 2:
Given I have received a password reset link
When I click the link after 24 hours
Then I see a message that the link has expired with an option to request a new one
AC 3:
Given I am resetting my password
When I enter a password shorter than 8 characters
Then I see a validation error and the password is not changed
```
### Tips
- Write 3-6 acceptance criteria per story — fewer means ambiguity, more means the story is too large
- Always include the sad path (error, edge case, timeout) not just the happy path
- Acceptance criteria are testable assertions — QA should be able to automate them directly
## Anti-Patterns to Avoid
### Too Large (Epic in Disguise)
- **Symptom**: "As a user, I want a complete reporting system"
- **Fix**: Split into individual reports, then split each report by filter/export/schedule
### Too Technical (Implementation Story)
- **Symptom**: "As a developer, I want to migrate the database to PostgreSQL"
- **Fix**: Reframe around user value: "As a user, I want search results in under 200ms" (which requires the migration)
### No Value Statement
- **Symptom**: "As a user, I want to click the submit button" — no "so that"
- **Fix**: Ask "why does the user care?" until you reach a meaningful benefit
### Compound Stories (AND in the Title)
- **Symptom**: "As a user, I want to search AND filter AND sort results"
- **Fix**: Split into three stories — search, filter, sort — each independently deliverable
### Solution-Prescriptive
- **Symptom**: "As a user, I want a dropdown menu with checkboxes"
- **Fix**: Describe the need, not the UI: "As a user, I want to select multiple categories to narrow results"
## Story Mapping
Organize stories into a two-dimensional map:
- **Horizontal axis**: User journey steps (left to right)
- **Vertical axis**: Priority (top = essential, bottom = nice-to-have)
Draw a horizontal line across the map to define your MVP — everything above the line ships first. This reveals gaps in the journey that individual story lists miss.
## Definition of Ready Checklist
- [ ] Story follows standard format with clear value statement
- [ ] INVEST criteria satisfied
- [ ] 3-6 acceptance criteria written in Given/When/Then
- [ ] Dependencies identified and resolved or accepted
- [ ] UX design reviewed (if UI-facing)
- [ ] Team has estimated the story
@@ -0,0 +1,95 @@
# Reviewing Artifacts (Product Lens)
When invoked as a reviewer, your role changes. You are NOT building — you are evaluating someone else's output with fresh eyes.
## Stance
- You did not produce this work. Judge the output, not the effort.
- You do not have access to the builder's reasoning (plan.md, memory.md). This is intentional — form independent judgment.
- Your job is to find gaps, ambiguities, and issues that would cause problems downstream.
- "READY" means a developer could implement from this without guessing. Not perfect — implementable.
## What to Check
### Requirements
- Is every requirement testable? (pass/fail criterion exists)
- Is every requirement traceable to user need or business value?
- Are there gaps? (things the intent implies but aren't covered)
- Are there contradictions?
- Are NFRs measurable? ("fast" → not measurable; "<200ms p95" → measurable)
- Is scope bounded? (what's explicitly out?)
### User Stories
- INVEST criteria met? (Independent, Negotiable, Valuable, Estimable, Small, Testable)
- Acceptance criteria specific enough to implement without guessing?
- Edge cases covered? (errors, empty states, boundaries)
- MVP boundary clear?
- Stories trace to requirements?
### Mockups/Wireframes
- All user stories have corresponding screens?
- Navigation flow complete? (every feature reachable)
- Error and empty states shown?
- Information hierarchy clear?
- Accessibility considered?
## How to Lodge Review Comments
Write your review to the review file the dispatch names (the `reviewFile` path
the request returned, under the intent record's `.aidlc-reviews/` directory).
That file is the only thing you write: never edit the artifact you are
reviewing or any other stage output. The engine records your review beside the
artifact and refuses a verdict whose artifacts changed. `ID` values are
stable (`R-01`, `R-02`, ...): never renumber, reuse, or change an existing ID.
`Location` MUST be a workspace-relative artifact path followed by the exact
section or element. `Required action` MUST state the concrete work in plain
language. On the first review, every finding has status `New`.
Use this exact format:
```markdown
## Review
**Verdict:** READY | NOT-READY
**Reviewer:** aidlc-product-lead-agent
**Date:** [ISO timestamp from Bash]
**Iteration:** [1, 2, etc.]
### Findings
| ID | Severity | Location | Finding | Required action | Status |
|---|---|---|---|---|---|
| R-01 | Critical | aidlc/spaces/<space>/intents/<intent-record>/inception/requirements-analysis/requirements.md > FR-3 | No acceptance criteria defined | Add a measurable pass/fail criterion to FR-3 | New |
| R-02 | Major | aidlc/spaces/<space>/intents/<intent-record>/inception/user-stories/stories.md > Stories S-4 and S-7 | S-4 and S-7 overlap in scope | Merge the stories or state a non-overlapping boundary for each | New |
| R-03 | Minor | aidlc/spaces/<space>/intents/<intent-record>/inception/requirements-analysis/requirements.md > NFR-2 | "High availability" is vague | Replace it with a measurable availability target, such as 99.9% | New |
### Summary
[1-2 sentences: overall assessment. What's the main issue holding it back, or why it's ready.]
```
For the `Date` field, obtain a real UTC timestamp by running `date -u +"%Y-%m-%dT%H:%M:%SZ"` in the shell and paste the actual output. Never guess or infer the date.
### Severity Levels
| Severity | Meaning | Blocks READY? |
|---|---|---|
| Critical | Cannot implement from this — fundamental gap or contradiction | Yes |
| Major | Implementable but will cause rework or confusion downstream | Yes (if >2 major findings) |
| Minor | Improvement opportunity, not blocking | No |
### Verdict Rules
- **READY** if: zero Critical, ≤2 Major (with clear workarounds), any number of Minor
- **NOT-READY** if: any Critical, OR >2 Major findings
### On Subsequent Iterations
When the dispatch brief includes `Prior findings (carry IDs forward)`:
- Treat that table as authoritative for prior human dispositions; it is
rendered from the audit ledger without rewriting the reviewed artifact.
- Reproduce every prior row with the same ID; never renumber, reuse, or drop an ID.
- Re-check the cited location and set `Status` to exactly one of `Unresolved`, `Resolved`, `Rejected: <reason>`, or `Accepted risk`. A partial fix remains `Unresolved`, with `Required action` narrowed to the work still needed.
- Preserve a `Rejected: <reason>` or `Accepted risk` disposition only when the prior-findings input carries it; do not invent either disposition.
- Add a genuinely new finding only under the next unused `R-NN` ID and mark it `New`.
- Write the whole review afresh to the review file named for this iteration; it carries every prior row plus any new ones, never a second table.
@@ -0,0 +1,78 @@
# NFR Reliability and Observability Guide
> This guide supplements the full NFR Requirements Guide held by the Security Engineer (lead agent for the NFR Requirements stage). It provides the QA Engineer with the reliability and observability sections needed for quality-focused contributions during NFR Requirements and Build and Test stages.
## Reliability Target-Setting
### SLA/SLO/SLI Hierarchy
| Term | Definition | Example | Who Defines |
|------|-----------|---------|-------------|
| **SLI** (Indicator) | The metric being measured | Successful requests / total requests | Engineering |
| **SLO** (Objective) | Internal target for the SLI | 99.95% success rate over 30 days | Engineering + Product |
| **SLA** (Agreement) | External contractual commitment | 99.9% uptime with financial penalties | Business + Legal |
**Rule**: SLO must be stricter than SLA. If SLA is 99.9%, set SLO at 99.95% to provide an internal buffer.
### Availability Targets and Their Implications
| Availability | Downtime / Year | Downtime / Month | Requires |
|-------------|-----------------|-------------------|----------|
| 99% (two nines) | 3.65 days | 7.3 hours | Basic monitoring, manual recovery |
| 99.9% (three nines) | 8.76 hours | 43.8 minutes | Auto-restart, health checks, alerting |
| 99.95% | 4.38 hours | 21.9 minutes | Multi-AZ, automated failover, load balancing |
| 99.99% (four nines) | 52.6 minutes | 4.38 minutes | Multi-region active-active, zero-downtime deploys |
| 99.999% (five nines) | 5.26 minutes | 26.3 seconds | Fully automated everything, extensive redundancy |
### Recovery Objectives
| Objective | Definition | How to Set |
|-----------|-----------|------------|
| **RTO** (Recovery Time Objective) | Maximum acceptable time to restore service | Based on business impact per hour of downtime |
| **RPO** (Recovery Point Objective) | Maximum acceptable data loss window | Based on cost of recreating or losing data |
| **MTTR** (Mean Time to Recovery) | Average time to restore from failure | Measured operationally; target should be < RTO |
| **MTBF** (Mean Time Between Failures) | Average time between failures | Measured operationally; drives reliability investment |
## Observability Requirements
### Three Pillars Specification
#### Metrics Requirements
Define for each component:
- **Business metrics**: Conversion rate, revenue processed, active users
- **Application metrics**: Request rate, error rate, latency (RED method)
- **Infrastructure metrics**: CPU, memory, disk, network (USE method)
- **Retention**: How long to keep each metric tier (1 min granularity for 7 days, 5 min for 30 days, 1 hour for 1 year)
#### Logging Requirements
| Log Level | When to Use | Retention | Indexing |
|-----------|------------|-----------|---------|
| ERROR | Unrecoverable failures requiring attention | 90 days | Full-text indexed |
| WARN | Recoverable issues, degraded behavior | 30 days | Full-text indexed |
| INFO | Significant business events, state transitions | 30 days | Structured fields only |
| DEBUG | Diagnostic detail for troubleshooting | 7 days | Not indexed (stored only) |
#### Tracing Requirements
- Distributed traces across all service boundaries
- Trace context propagation via W3C Trace Context headers
- Sampling strategy: 100% for errors, 10% for normal traffic (adjust for volume)
- Trace retention: 7 days at full detail, 30 days for trace metadata
### Alerting Requirements Template
For each alert:
```
Alert: [descriptive name]
SLI: [which service level indicator]
Threshold: [when to fire -- e.g., error rate > 1% for 5 minutes]
Severity: [page (wake someone up) | ticket (next business day) | log (informational)]
Runbook: [link to response procedure]
Notification: [who gets notified via which channel]
Auto-remediation: [if any automated response is triggered]
```
### Observability Anti-Patterns to Avoid
- Alerting on causes instead of symptoms (alert on error rate, not CPU usage)
- Missing correlation IDs across service boundaries
- Logging sensitive data (PII, credentials, tokens)
- Alert fatigue from noisy thresholds (tune before deploying)
- Dashboard sprawl without clear ownership (every dashboard needs an owner)
@@ -0,0 +1,84 @@
# Non-Functional Requirement (NFR) Validation Methods
Practical approaches to validating performance, scalability, and reliability requirements.
## Load Testing Tools and Methodology
**Tools**:
- **k6** (Grafana): Script-based, developer-friendly, runs locally or in cloud. Preferred for API load testing.
- **Locust** (Python): Distributed, programmable load generation. Good for complex user behaviour simulation.
- **Artillery**: YAML-driven, supports HTTP/WebSocket/Socket.io. Quick to set up.
- **AWS Distributed Load Testing**: CloudFormation-based, uses Fargate to generate load from within AWS.
**Methodology**:
1. Identify critical user journeys and their expected traffic volumes.
2. Create realistic test scripts with think times, parameterized data, and varied payloads.
3. Establish a performance baseline on current production or staging.
4. Run tests against an environment that mirrors production (same instance sizes, data volume, network topology).
5. Collect results, compare against NFR targets, identify bottlenecks.
## Performance Test Design Patterns
### Ramp-Up Test
Gradually increase virtual users from 0 to target over 5-15 minutes. Validates system behaviour under increasing load and identifies the breaking point.
### Steady-State Test
Hold constant load at expected peak for 30-60 minutes. Validates sustained performance, memory leaks, connection pool exhaustion, and resource saturation.
### Spike Test
Suddenly inject 3-5x normal load for a short burst (2-5 minutes). Validates auto-scaling triggers, queue depth handling, circuit breaker behaviour, and graceful degradation.
### Soak Test
Run at moderate load (60-80% of peak) for 4-24 hours. Detects slow memory leaks, file handle exhaustion, log rotation issues, and gradual performance degradation.
## Latency Percentiles
- **p50 (median)**: The typical user experience. Target depends on use case (web API: < 100ms).
- **p95**: Captures most users' worst-case experience. The primary SLO metric for most services.
- **p99**: Tail latency. High p99 with low p50 indicates inconsistent performance (GC pauses, cold starts, noisy neighbours).
- **p99.9**: Relevant for high-volume services where even 0.1% affects thousands of requests.
Always measure percentiles, not averages. An average of 50ms can hide a p99 of 5 seconds.
## Throughput Measurement
- Measure in requests per second (RPS) for APIs, messages per second for queues, transactions per second (TPS) for databases.
- Record throughput alongside latency; throughput without latency context is meaningless.
- Identify the throughput ceiling: the RPS at which latency begins to degrade beyond acceptable thresholds.
- For batch processing, measure records processed per second and total job duration.
## Capacity Planning
1. Determine current peak traffic from production metrics (CloudWatch, X-Ray).
2. Apply a growth multiplier (typically 2-3x current peak for 12-month horizon).
3. Load test at the projected peak to validate infrastructure can handle it.
4. Identify the scaling bottleneck (database connections, Lambda concurrency, NAT gateway bandwidth).
5. Document the cost at projected peak for budget approval.
## Auto-Scaling Validation
- Test that scale-out triggers fire within acceptable time (target: under 2 minutes for EC2, under 30 seconds for Lambda).
- Test that scale-in does not cause request failures (connection draining, graceful shutdown).
- Validate scaling policies use the right metric (CPU alone is often insufficient; include request count, queue depth).
- Test minimum and maximum capacity limits to prevent runaway scaling costs.
- Simulate a scaling event during a spike test and measure the latency impact during scale-out.
## NFR Target-vs-Actual Matrix
Track every NFR with a structured comparison:
| NFR | Target | Actual | Status | Test Date | Notes |
|-----|--------|--------|--------|-----------|-------|
| API latency p95 | < 200ms | 145ms | PASS | 2024-01-15 | Under 500 RPS |
| API latency p99 | < 500ms | 620ms | FAIL | 2024-01-15 | DB connection pool saturation |
| Throughput | > 1000 RPS | 1250 RPS | PASS | 2024-01-15 | |
| Availability | 99.95% | — | PENDING | — | Requires 30-day measurement |
Review the matrix at each milestone. Failing NFRs are risks that must be addressed before release.
## SLA Testing
- Simulate failure scenarios (AZ failure, dependency timeout, database failover) and measure recovery time.
- Validate that SLA commitments (uptime percentage, response time) hold under degraded conditions.
- Test circuit breakers, retries, and fallback responses under dependency failure.
- Document the gap between internal SLOs (tighter) and external SLAs (looser) to provide a safety margin.
@@ -0,0 +1,73 @@
# Test Strategy Patterns
Comprehensive guidance for building a test strategy that balances confidence, speed, and maintainability.
## The Test Pyramid
Structure tests in layers, with volume decreasing as scope and cost increase:
1. **Unit Tests (base, ~70%)** — Test individual functions, classes, or modules in isolation. Fast, deterministic, run on every commit. Mock external dependencies.
2. **Integration Tests (middle, ~20%)** — Test interactions between components: service-to-database, service-to-service, message producer-to-consumer. Use real dependencies where practical (testcontainers, localstack).
3. **End-to-End Tests (top, ~10%)** — Test complete user workflows through the full stack. Slowest and most brittle; keep the count small and focused on critical paths.
Anti-pattern: the **ice cream cone** (mostly manual/e2e tests, few unit tests) leads to slow feedback and flaky pipelines.
## Test Doubles
| Type | Purpose | Example |
|------|---------|---------|
| **Mock** | Verify interactions (was method X called with args Y?) | `jest.fn()`, `unittest.mock.Mock` |
| **Stub** | Return predetermined data; no interaction verification | Hard-coded return values |
| **Fake** | Working implementation with shortcuts (in-memory DB) | SQLite for integration tests |
| **Spy** | Wraps real object; records calls while executing real logic | `jest.spyOn()`, Sinon spies |
Guidelines:
- Prefer stubs over mocks to keep tests less coupled to implementation.
- Use fakes (LocalStack, testcontainers) for integration tests to increase realism.
- Avoid mocking what you do not own; wrap third-party libraries behind an interface and mock the interface.
## Test Data Management
- Use **factories** or **builders** to construct test data programmatically (factory_boy, fishery, @faker-js/faker).
- Each test should set up its own data; avoid shared mutable state across tests.
- For integration tests, use database transactions that roll back after each test, or truncate tables in setup.
- Maintain a **seed data** script for local development and CI that populates reference data consistently.
- Mask or synthesize PII for test environments; never copy production data without anonymization.
## Contract Testing
- Use Pact or similar tools to verify that a consumer's expectations match the provider's actual API.
- Consumer writes a contract (expected request/response pairs); provider verifies it independently.
- Decouple consumer and provider deployments: each team runs contract tests in their own pipeline.
- Essential for microservice architectures where integration tests across all services are impractical.
## Property-Based Testing
- Instead of specific input/output examples, define properties that must hold for all valid inputs.
- Tools: Hypothesis (Python), fast-check (TypeScript), QuickCheck (Haskell-inspired).
- Example properties: "serializing then deserializing returns the original value", "sorting is idempotent", "output list length equals input list length".
- Excellent for finding edge cases that example-based tests miss (empty strings, negative numbers, Unicode).
## Mutation Testing
- Introduces small changes (mutations) to source code and checks whether tests catch them.
- A surviving mutant means a gap in test assertions.
- Tools: Stryker (JS/TS), mutmut (Python), pitest (Java).
- Use selectively on critical modules; full-codebase mutation testing is expensive.
## Coverage Metrics and Targets
- **Line coverage**: Percentage of lines executed. Baseline target: 80%.
- **Branch coverage**: Percentage of conditional branches taken. More meaningful than line coverage.
- **Mutation score**: Percentage of mutants killed. Gold standard but costly.
- Coverage is a necessary but not sufficient quality signal. 100% coverage with weak assertions catches nothing.
- Enforce coverage in CI as a gate: fail the build if coverage drops below the threshold.
## CI Integration
- Run unit tests on every push; gate merges on green status.
- Run integration tests on pull request creation and nightly.
- Run e2e tests nightly or on release branches; do not block every PR.
- Parallelize test suites across CI workers to keep feedback under 10 minutes.
- Report test results as PR annotations (JUnit XML, GitHub Actions test reporter).
- Track flaky tests explicitly; quarantine or fix within one sprint.
@@ -0,0 +1,132 @@
# Testing Guide
## Test Pyramid
Structure tests in this ratio (approximate):
```
/ E2E \ ~5% (slow, expensive, high confidence)
/ Integration \ ~20% (moderate speed, cross-component)
/ Unit \ ~75% (fast, isolated, high volume)
```
### Unit Tests
- Test a single function/method/class in isolation
- Mock all external dependencies (database, network, file system)
- Execute in milliseconds; run on every commit
- Target: every public function with non-trivial logic
- Naming: `test_[function]_[scenario]_[expected_result]`
### Integration Tests
- Test interactions between two or more components
- Use real dependencies where practical (test database, local queue)
- Execute in seconds; run on every PR
- Target: API endpoints, service-to-repository interactions, message flows
- Focus on contract verification between components
### End-to-End Tests
- Test complete user workflows through the full stack
- Use a deployed (or containerized) environment
- Execute in minutes; run before release
- Target: critical business workflows only (login, core CRUD, payment)
- Maximum 20-30 e2e tests; more indicates missing integration coverage
## Test Types Beyond the Pyramid
### Performance Tests
- **Load test**: Expected concurrent users for sustained period (establish baseline)
- **Stress test**: Increase load until failure (find breaking point)
- **Soak test**: Sustained load over hours (find memory leaks, connection exhaustion)
- Metrics: response time (p50, p95, p99), throughput, error rate, resource utilization
### Security Tests
- **SAST**: Static analysis of source code for known vulnerability patterns
- **DAST**: Dynamic testing of running application (injection, XSS, auth bypass)
- **Dependency scan**: Check third-party packages against CVE databases
- Integrate into CI pipeline as blocking quality gate
### Contract Tests
- Verify API consumer expectations match provider implementation
- Use Pact or similar consumer-driven contract framework
- Run independently of full integration environment
- Essential for microservices and multi-team projects
### Accessibility Tests
- Automated: axe-core or pa11y for WCAG violations
- Manual: keyboard navigation, screen reader walkthrough
- Run automated checks in CI; manual checks before release
## Depth-Aware Test Volume
Test volume scales with the active test strategy (defaults to depth level, overridable via `--test-strategy`). The pyramid ratios above set proportions within the volume cap for each level:
| Strategy | Tests per Component | Test Types | Total (typical) |
|----------|-------------------|------------|-----------------|
| Minimal (Nyquist) | 1 per requirement + happy-path floor | Unit only | ~5-15 |
| Standard | 5-8 | Unit + integration | ~20-50 |
| Comprehensive | 10-15 | Unit + integration + E2E + perf + security | ~50-100+ |
- **Minimal** uses a requirement-driven model (1 test per requirement, not per component). The pyramid doesn't apply — unit tests only.
- **Standard** and **Comprehensive** use a per-component model. The pyramid proportions (75/20/5) apply within the generated set.
- All levels are soft guidelines. The LLM can exceed when context demands (e.g., security-critical code at Minimal depth).
See stage-protocol.md §8 "Test Strategy" for the authoritative guidance.
## Test Case Design Template
```
Test ID: TC-[NNN]
Requirement: [FR/NFR ID]
Title: [descriptive name]
Priority: [P0/P1/P2]
Type: [unit/integration/e2e/performance/security]
Preconditions: [setup required]
Steps:
1. [action]
2. [action]
Expected Result: [verifiable outcome]
Cleanup: [teardown required]
```
## Quality Gate Definitions
### Gate 1: Code Review (before merge)
- All unit tests pass
- No new linter warnings
- Code coverage does not decrease
- Security scan finds no high/critical issues
### Gate 2: Integration (after merge to main)
- All integration tests pass
- Contract tests pass
- Performance baseline not regressed (>10% degradation = fail)
### Gate 3: Release Readiness (before deploy)
- All e2e tests pass on staging environment
- Accessibility automated checks pass
- No open P0/P1 defects
- Stakeholder acceptance sign-off obtained
## Test Data Strategy
- **Factories**: Generate test objects with sensible defaults, override per test
- **Fixtures**: Static data loaded before test suite (reference data, lookup tables)
- **Builders**: Fluent API for constructing complex test scenarios
- **Synthetic**: Generated data for performance tests (realistic volume and distribution)
- **Isolation**: Each test owns its data. Never share mutable state between tests.
- **Cleanup**: Tests clean up after themselves. Use database transactions that roll back.
## Defect Report Format
```
Defect ID: BUG-[NNN]
Severity: [Critical/High/Medium/Low]
Component: [affected component]
Summary: [one-line description]
Reproduction Steps:
1. [step]
2. [step]
Expected: [what should happen]
Actual: [what actually happens]
Environment: [OS, browser, version, config]
Evidence: [logs, screenshots, error messages]
```
@@ -0,0 +1,35 @@
# AI-DLC Methodology Principles
## Design Principle: Small Mob, Broad Agents
AI-DLC is built on the mob model — a small cross-functional group moving fast together. The agents mirror that. Rather than dozens of narrow specialists (which recreates waterfall handoff chains), we define **11 broadly capable agents** that each participate across multiple stages and phases, just as a real architect or developer would in a mob session.
Each agent carries context across stages because they are present throughout. This eliminates handoffs, reduces coordination overhead, and keeps the process agile.
## Core Principles
1. **User decides, AI executes** — Every material decision goes through an approval gate where the user reviews, revises, or overrides.
2. **Adaptive depth** — Simple projects skip heavyweight stages. Complex projects get full coverage. The workflow adapts to project needs.
3. **Traceable artifacts** — Every stage produces versioned markdown documents in `aidlc-docs/`, creating a complete decision record.
4. **Multi-role expertise** — Each stage is guided by domain-expert agent personas to ensure appropriate depth.
5. **No emergent behavior** — Agents follow prescribed protocols. Approval menus, completion messages, and state transitions are standardized.
6. **Questions before assumptions** — When in doubt, ask. Incomplete answers lead to poor designs.
7. **Contradiction detection** — Cross-check all answers for scope mismatches, risk mismatches, and technology conflicts.
## Five-Phase Structure
| Phase | Purpose | Key Outcome |
|-------|---------|-------------|
| **INITIALIZATION** | Bootstrap — state files, directory scaffold, workspace scan, routing | Configured workspace ready for workflow |
| **IDEATION** | Validate the initiative — intent, market, feasibility, scope, team | Approved initiative brief |
| **INCEPTION** | Elaborate — requirements, stories, design, architecture, units, delivery plan | Detailed execution plan |
| **CONSTRUCTION** | Build — functional design, NFRs, infrastructure, code, tests, CI | Working tested code |
| **OPERATION** | Deploy & operate — pipelines, environments, observability, incidents, feedback | Production system with monitoring |
## Scope System
Not every task requires every stage. Scopes (see the compiled scope grid or run `/aidlc --doctor` for the enabled set) determine which stages execute and at what depth.
## Self-Learning Guardrails
When a human corrects agent behavior, the correction becomes a permanent guardrail so the mistake never repeats. Guardrails are classified as organization-level (all projects) or project-level (this repo only).
@@ -0,0 +1,396 @@
# Audit Event Taxonomy
**Event names MUST match this table exactly.** Do not invent new event types. For stage completions, ALWAYS use `STAGE_COMPLETED` — do not substitute stage-specific names like "Requirements Analysis Complete" or "Code Generated".
Stage lifecycle is report-owned: the conductor requests gate, revision,
completion, and skip outcomes through `tools/aidlc-orchestrate.ts report`.
The state tool entries below are the engine's internal atomic emitters, not
commands a stage or conductor invokes directly.
> See [`docs/reference/12-state-machine.md`](../../../../docs/reference/12-state-machine.md) for the state transitions that emit each event. Events marked `✓` are MANDATORY and asserted by `tests/feature/t48-audit-event-emitters.sh`.
## Naming Convention
All event names follow `SUBJECT_PAST_VERB` — every event answers "what happened?"
## Emitter-Owned Fields
The structured renderer writes exactly one `Timestamp` and one `Event` line per
block; callers must not supply either field. For compatibility, the generic
`audit append --field Timestamp=...` form is still accepted, but its value is
intentionally ignored. Historical shards are not rewritten: readers that parse
whole files must split on `---` and use the first timestamp in each block, or
deduplicate timestamp fields produced by older versions.
## Event Registry (95 events, 23 categories)
### Workflow Lifecycle (4 events)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| ✓ `WORKFLOW_STARTED` | Scope determined, workflow begins | Timestamp, Scope, Request; optional Source Baseline (`sha256:<listing-hash>` or `unbindable`) | `tools/aidlc-utility.ts intent-create` |
| ✓ `WORKFLOW_COMPLETED` | All in-scope stages done | Timestamp, Scope, Details | `tools/aidlc-state.ts complete-workflow` |
| ✓ `WORKFLOW_PARKED` | Workflow parked mid-flow for a later session (no stage advanced) | Timestamp, Stage | `tools/aidlc-state.ts park` |
| ✓ `WORKFLOW_UNPARKED` | Park marker cleared on explicit `--resume` re-entry | Timestamp | `tools/aidlc-state.ts unpark` |
### Phase Lifecycle (4 events)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| ✓ `PHASE_STARTED` | Phase begins (first in-scope stage about to run) | Timestamp, Phase, Stage count, Scope | `tools/aidlc-utility.ts intent-create` (Init phase), `tools/aidlc-state.ts advance` (phase boundary) |
| ✓ `PHASE_COMPLETED` | Crossed a phase boundary | Timestamp, From phase, To phase, Stages completed | `tools/aidlc-state.ts advance`, `tools/aidlc-state.ts complete-workflow` |
| `PHASE_VERIFIED` | Traceability check at boundary | Timestamp, Phase boundary, Pass/fail, Issues | `tools/aidlc-state.ts advance`, `tools/aidlc-state.ts complete-workflow` |
| `PHASE_SKIPPED` | Scope excludes phase | Timestamp, Phase, Scope, Reason | `tools/aidlc-utility.ts intent-create` (per-phase scope eval) |
### Stage Lifecycle (6 events)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| ✓ `STAGE_STARTED` | Stage enters `[-]` Active | Timestamp, Stage, Agent; optional Source Baseline when the entered stage declares `workspace_requires` | `tools/aidlc-state.ts advance`, `tools/aidlc-utility.ts intent-create` (init stages) |
| `STAGE_AWAITING_APPROVAL` | Stage enters `[?]` (gate open) | Timestamp, Stage, Artifacts, optional `Recovered=true` (backfilled gate row) or `Revalidated=true` (an already-open gate consumed a blocking-sensor override); a blocking-sensor override also carries Blocking Sensor Override, Blocking Sensor IDs, optional Blocking Sensor Detail Paths, and Blocking Sensor Reasons. Its authorization is the preceding exact `DECISION_RECORDED` → `HUMAN_TURN` → `QUESTION_ANSWERED` pair. | `tools/aidlc-state.ts gate-start` (organic, `--recovered` backfill, or override revalidation), `tools/aidlc-state.ts revise` (gate re-entry), `tools/aidlc-state.ts approve` (backstop re-entry after its opening guards pass) |
| `STAGE_REVISING` | Stage enters `[R]` (user rejected gate) | Timestamp, Stage, Revision count, Feedback, optional `Recovered=true` (backfilled by the approve-time revision backstop) | `tools/aidlc-state.ts reject`, `tools/aidlc-state.ts approve` (backstop backfill) |
| ✓ `STAGE_COMPLETED` | Stage finishes (`[x]`) | Timestamp, Stage, Details, Artifacts | `tools/aidlc-state.ts approve` (gated stages; also auto-advances to next), `tools/aidlc-state.ts advance` (non-gated stages), `tools/aidlc-utility.ts intent-create` (init stages) |
| `STAGE_JUMPED` | Forward/backward/redo jump target reached | Timestamp, Direction, Source, Target, Scope; optional Source Baseline. Backward jumps also carry JSON arrays for Changed Upstream Artifacts, Invalidated Downstream Artifacts, and Invalidated Downstream Reviews | `tools/aidlc-jump.ts execute` |
| `STAGE_SKIPPED` | Current stage reports a justified skip, or a jump skips it (`[S]`) | Timestamp, Stage, Reason | `tools/aidlc-state.ts skip` (internally routed by `aidlc-orchestrate.ts report --result skipped`), `tools/aidlc-jump.ts execute` |
### Session Events (5 events — hook-owned, independent of workflow lifecycle)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `SESSION_STARTED` | Fresh Claude Code session begins (source=startup or clear) | Timestamp, Source | `hooks/aidlc-session-start.ts` |
| `SESSION_RESUMED` | Existing Claude Code session resumed (source=resume) | Timestamp, Source | `hooks/aidlc-session-start.ts` |
| `SESSION_COMPACTED` | Context compaction occurred | Timestamp, Current Stage, State Validity | `hooks/aidlc-validate-state.ts` (PreCompact) |
| `SESSION_ENDED` | Claude Code session terminates | Timestamp, Reason | `hooks/aidlc-session-end.ts` |
| `HUMAN_TURN` | A supported prompt-submit or answered-widget seam was observed (the approval/interview gate requires one since the last gate resolution); omitted when the driver declares `AIDLC_UNATTENDED=1` | Timestamp | `hooks/aidlc-record-human-turn.ts` (UserPromptSubmit + PostToolUse AskUserQuestion) + the per-harness prompt-submit adapters |
`HUMAN_TURN` is chronological presence evidence, not authenticated decision
content. The `--user-input`, `--feedback`, and `--details` fields recorded by
authority-bearing tools are caller-supplied prose. A narrow defense-in-depth
tripwire rejects recognized explicit conductor/model self-attribution, but
unlabelled, paraphrased, localized, or otherwise unrecognized wording remains
indistinguishable from human-authored text. `AIDLC_SKIP_HUMAN_PRESENCE_GUARD=1`
also disables that tripwire for deterministic recovery/tests. Audit shards are
operational evidence, not a tamper-proof human-authorship boundary.
### Initialization Events (3 events — fire IN ADDITION TO `STAGE_COMPLETED`)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `WORKSPACE_SCAFFOLDED` | Directory tree created | Timestamp, Details | `tools/aidlc-utility.ts` handleInit |
| `WORKSPACE_SCANNED` | Workspace detection done | Timestamp, Project type, Details | `tools/aidlc-utility.ts` handleInit |
| `WORKSPACE_INITIALISED` | State file created | Timestamp, Details | `tools/aidlc-utility.ts` handleInit |
### Navigation Events (7 events)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `SCOPE_CHANGED` | `--scope` changed existing scope | Timestamp, Old scope, New scope | `tools/aidlc-utility.ts` |
| `PLUGIN_SELECTION_CHANGED` | `select-plugins` changed enabled plugins | Timestamp, Previous Selection, New Selection | `tools/aidlc-utility.ts select-plugins` |
| `DEPTH_CHANGED` | `--depth` changed depth level | Timestamp, Old depth, New depth | `tools/aidlc-utility.ts` |
| `TEST_STRATEGY_CHANGED` | `--test-strategy` changed test strategy | Timestamp, Old strategy, New strategy | `tools/aidlc-utility.ts` |
| `REVIEW_CLASS_CHANGED` | `--review` changed the per-run review override | Timestamp, Old Override, New Override | `tools/aidlc-utility.ts` |
| `SCOPE_DETECTED` | Auto-detected from freeform text | Timestamp, Detected scope, Input text, Source, Matched keywords (optional; present when `Source=keyword`) | `tools/aidlc-utility.ts detect-scope` |
| `RECOMPOSED` | The adaptive composer re-shaped a running workflow's pending stages (suffix flips via `recompose`) | Timestamp, Scope, Stages skipped, Stages added, Stages in Scope | `tools/aidlc-utility.ts recompose` |
### Change Control Events (2 events)
Change Control is one per-intent setting, `strict` or `relaxed`, that decides what a governed checkpoint does when an input changed after the human approved or confirmed something. Both rows are provenance emitted through the library by their owners; the public audit CLI refuses them.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `CHANGE_CONTROL_SET` | The intent's Change Control value moved: the `change-control <strict\|relaxed>` verb (from the `--change-control` flag or a plain-chat request) rewrote the state line, a `scope-change` carried a scope-supplied value to the new scope's default, or a governed checkpoint observed that a memory layer edit changed the effective value for this running intent | Timestamp, Old Value, New Value, Source (`you`, `scope <name>`, or `<layer>.md`) | `tools/aidlc-lib.ts` (`recordChangeControlSet` for the verb and `scope-change`, `governedChangeControl` for the checkpoints) |
| `CHANGE_ACCEPTED` | A governed checkpoint found that an input changed after a human approval or confirmation and, under `relaxed`, recorded the change and continued instead of refusing. Written once per distinct change: the same Recorded and Current values for the same Checkpoint, Stage, and Unit never produce a second row | Timestamp, Stage, optional Unit, Checkpoint (`plan-approval`, `review-receipt`, `summary-confirmation`), Changed (a bounded path list or `(paths unavailable)`), Recorded, Current, Details (the one line the human hears) | `tools/aidlc-lib.ts` (`recordAcceptedChanges`, called by the checkpoint owners: `aidlc-log.ts decision` / `answer` / `review`, `aidlc-testing-posture.ts begin`, `aidlc-state.ts` gate and completion checks, the plan-approval guard hook) |
### Interaction Events (11 events)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `DECISION_RECORDED` | Before presenting a non-gate structured question, to record the options shown. Consolidated-summary prompts also carry checkpoint identity | Timestamp, Stage, Decision, Options; optional Checkpoint, Questions File, Unit, Attempt Generation, Workflow | `tools/aidlc-log.ts decision` |
| `GATE_APPROVED` | Human approved at gate | Timestamp, Stage, User Input; optional Review Finding Dispositions (versioned JSON mapping every current New/Unresolved review finding to Accepted risk, keyed by review artifact, finding ID, and finding-content fingerprint), Unit, Gate Scope, Gate Stages, Attempt Generation (team Unit gates); Unit merge gates also carry Pinned OID, Strategy, Target branch | `tools/aidlc-state.ts approve`, `tools/aidlc-unit.ts gate` |
| `GATE_REJECTED` | Human requested changes | Timestamp, Stage, Feedback; optional Review Finding Dispositions (versioned JSON for findings the human explicitly rejected with an exact reason), `Recovered=true` (backfilled by the approve-time revision backstop), Prior Accepted Source Fingerprint (the prior attempt's validated final swarm aggregate; never a replacement completion baseline), Unit, Gate Scope, Gate Stages, Attempt Generation (team Unit gates); Unit merge gates also carry Pinned OID, Strategy, Target branch | `tools/aidlc-state.ts reject`, `tools/aidlc-state.ts approve` (backstop backfill), `tools/aidlc-unit.ts gate` |
| `QUESTION_ANSWERED` | Non-gate question answered by user | Timestamp, Stage, Details; optional Unit, Attempt Generation | `tools/aidlc-log.ts answer` |
| `SUMMARY_CONFIRMATION_RECORDED` | Consolidated-summary choice recorded after the matching prompt and a fresh human turn; reserved from the public audit CLI | Timestamp, Stage, Details, Checkpoint, Questions File, Questions SHA-256, Hash Scope (required on new receipts; legacy rows may omit it), Summary Authorization Id (on a `Looks correct` receipt: sha256 over the attempt, stage, Unit, workflow, questions path, confirmed-content hash, and choice; identical confirmations mint the same id, changed answers a new one; legacy rows omit it); optional Unit, Workflow | `tools/aidlc-log.ts answer --checkpoint summary-confirmation` |
| `PLAN_APPROVAL_RECORDED` | Provenance that Code Generation consumed a protected Plan Approval challenge/response receipt; this Markdown row is not authorization evidence. `Directive Epoch` is recorded for diagnosis and never compared: the approval binds to `Plan Target`, `Run floor` (the stage attempt), and `Approval Fingerprint` (the plan content) | Timestamp, Stage, Details, Checkpoint, Plan Target, Intent, Directive Epoch, Run floor, Approval Fingerprint, Questions File, Questions SHA-256, Prompt SHA-256, Session | `tools/aidlc-log.ts answer --checkpoint plan-approval` |
| `PLAN_APPROVAL_OVERRIDDEN` | The human broke the glass: they typed `Override Plan Approval: <reason>` as a prompt (recorded by the human-turn hook, never from a picked option) and the conductor ran `answer --override` with that reason after the normal receipt path refused. The receipt this row accompanies binds to plan content and stage attempt only; source certification is skipped for it. Always paired with a `PLAN_APPROVAL_RECORDED` row carrying `Override: yes`; reserved from the public audit CLI | Timestamp, Stage, Reason (verbatim), Failed Checks (the normal path's refusals, separated by a vertical bar), Session, Unit (or `stage-level`), Fingerprint | `tools/aidlc-log.ts answer --checkpoint plan-approval --override` |
| `REVIEW_REQUESTED` | Conductor dispatches the §12a reviewer sub-agent; reserved from the public audit CLI. Structurally malformed rows are non-authoritative and do not consume an ordinal | Timestamp, Stage, Reviewer, Iteration, Artifact Fingerprint (`sha256:<hex>` from one stable snapshot of every declared artifact exactly as dispatched for review), Request Id (`review:<32 lowercase hex>`, minted per request; the completion row and the review record echo it), optional Unit + Attempt Generation (authoritative-DAG per-unit claims), Source Fingerprint on `workspace_requires` stages, Unit Source Fingerprint (manifest bytes + claimed source listing) on per-unit `workspace_requires` stages, optional Retry (`pending-request`, accepted once after a modern binding exists), optional Upgrade (`legacy-request`, one bounded modernization of a valid request recorded before request ids or source binding), optional Recovery (`stale-receipt`) plus Recovery Cause (`artifact`, `source`, or `artifact+source`). Rows written under the retired appendix protocol also carry Review Appendix Artifact, Review Appendix Offset, Review Appendix Prior Digest, Review Appendix Prior Length, and Review Challenge; they stay readable, and a completion of such a request echoes them unchanged so the pair still binds | `tools/aidlc-log.ts review` |
| `REVIEW_COMPLETED` | Reviewer verdict recorded; gates approval and is reserved from the public audit CLI. Malformed rows are ignored without consuming their pending request | Timestamp, Stage, Reviewer, Iteration, Verdict, Request Fingerprint (must match the request), Artifact Fingerprint (the same stable snapshot: the reviewer writes no artifact, so the reviewed bytes are the requested bytes), Request Id (must match the request), Review Record (record-relative path `.aidlc-reviews/<stage>/stage/<attempt>/<iteration>.json` or `.aidlc-reviews/<stage>/units/<unit>/<attempt>/<iteration>.json`) + Review Record Digest (`sha256:<hex>` over the record bytes; a record that no longer hashes to it is not the review), optional Unit + Attempt Generation (per-unit claims), Request Source Fingerprint + Source Fingerprint (identical request-time source identity on `workspace_requires` stages), Unit Source Fingerprint or Unit Source Binding Bypass on per-unit `workspace_requires` stages, Review Challenge only for the deprecated appendix migration path, Recovery only for the one stale-receipt recovery pass, optional Workflow for isolated runs | `tools/aidlc-log.ts review --verdict` |
| `PIPELINE_LINK_COMPLETED` | A declared pipeline link returned in order; current-attempt receipts gate pipeline approval and are reserved from the public audit CLI | Timestamp, Stage, Link, Position (`k/N`), optional Repo (required by the protocol for multi-repo chains), optional Workflow (`single-stage:<slug>` for isolated runs) | `tools/aidlc-log.ts link` |
`Hash Scope: confirmed-content-v1` identifies the semantic questions-file digest
used by newly emitted receipts. It normalizes line endings, preserves the
original order of the preamble and confirmed sections, and trims trailing
whitespace from the resulting canonical content. It includes every visible
Q<n> section and each `Requested Changes Feedback` section, including follow-up
questions added after an assumption decision. Exactly one visible top-level
`Assumption Confirmation` section is valid only after the summary and is
excluded, along with its contents; a same-named pre-summary section remains part
of the confirmed digest. The excluded section's assumptions and answer are not
covered by the digest and remain subject to the stage's existing decision/answer
and sensor checks. Any other visible Markdown or
raw-HTML heading after the summary is invalid. Heading-like text in HTML
comments, code spans, fenced or indented code, and HTML attribute values is not
a section.
A receipt with no `Hash Scope` retains the legacy whole-file digest contract;
an in-flight legacy receipt therefore needs a fresh human confirmation to
create a scoped receipt before an allowed post-confirmation append can recover.
Any other scope is rejected.
Summary-confirmation authority comparisons preserve append order within one
audit shard. Across shards, different timestamps establish order; equal
timestamps are causally unordered. Completion fails closed when such a tie
could change the current attempt, selected receipt, or whether an artifact was
written after confirmation, and requires fresh evidence with a later timestamp.
### Unit Configuration and Lifecycle Events (7 events — unit-major Construction)
The interactive twin of the swarm's `SWARM_UNIT_*` ledger. `UNIT_COMPLETED` is
the completion receipt the engine's coverage walk prefers over bare artifact
existence once any receipt exists for the stage; the emitting verb verifies
the unit's required artifacts as regular files on disk before committing it.
`Run floor` is an exact boundary token (`<event>:<timestamp>#<ordinal>`), so
same-second attempts within one shard cannot reuse receipts. Equal-time
boundaries in different shards are causally unordered and use a deterministic
`AMBIGUOUS:<timestamp>#<digest>` floor; prior receipts cannot match it.
Unit-major stages key the floor to workflow/jump/rejection boundaries because
their work can precede their own `STAGE_STARTED`.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `UNIT_OWNERSHIP_SET` | Unit-major ownership mode is set before unit activity starts | Timestamp, Mode | `tools/aidlc-state.ts set-unit-ownership` |
| `UNIT_GATE_RHYTHM_SET` | Team-owned gate rhythm is set before unit activity starts | Timestamp, Rhythm | `tools/aidlc-state.ts set-unit-gate-rhythm` |
| `UNIT_STARTED` | A unit's work begins on an inline per-unit stage; refused while another unit of the stage is open | Timestamp, Stage, Unit, Run floor; optional Attempt Generation | `tools/aidlc-state.ts unit start` |
| `UNIT_PAUSED` | A unit stops before completion; the checkpoint carries why and what comes next | Timestamp, Stage, Unit, Run floor, Reason, Next Action; optional Attempt Generation | `tools/aidlc-state.ts unit pause` |
| `UNIT_RESUMED` | The paused unit is explicitly resumed (the engine hard-stops until this) | Timestamp, Stage, Unit, Run floor; optional Attempt Generation | `tools/aidlc-state.ts unit resume` |
| `UNIT_COMPLETED` | The unit's work is done AND its required artifacts are regular files on disk (verified at emit) | Timestamp, Stage, Unit, Run floor; optional Attempt Generation | `tools/aidlc-state.ts unit complete` |
| `UNIT_MERGED` | Main landed the pinned candidate content and folded this Unit's row; transported receipts now satisfy main's floors | Timestamp, Unit, Owner, Pinned OID, Merge commit OID, Attempt Generation | `tools/aidlc-state.ts fold-unit-merge` |
### Artifact Events (3 events — hook-emitted)
The artifact hook emits for writes in either the active intent's record tree or
the active space's shared `codekb/<repo>/` tree.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `ARTIFACT_CREATED` | New artifact written in the active intent record or space-level codekb tree | Timestamp, Tool, File, Context, optional Summary Authorization Id (the active summary confirmation for the written stage and Unit at write time, read from `<record>/.aidlc-summary-authorization/<stage>/stage.json` or `<record>/.aidlc-summary-authorization/<stage>/units/<unit>.json`; absent when no confirmation is active) | `hooks/aidlc-write-audit-log.ts` (PostToolUse; Write to net-new path) |
| `ARTIFACT_UPDATED` | Existing artifact modified in either tree | Timestamp, Tool, File, Context, optional Summary Authorization Id (as for `ARTIFACT_CREATED`) | `hooks/aidlc-write-audit-log.ts` (PostToolUse; Edit, or Write overwriting existing) |
| `ARTIFACT_REUSED` | Re-use decision on backward jump or per-repo pipeline reuse evidence; only `Decision=keep` grants the pipeline exemption; reserved from the public audit CLI | Timestamp, Stage, Decision, Artifacts, optional Repo, optional Workflow (`single-stage:<slug>` for isolated freshness-bound reuse) | `tools/aidlc-state.ts reuse-artifact` |
### Subagent Events (1 event — hook-emitted)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `SUBAGENT_COMPLETED` | Subagent task finishes | Timestamp, Agent Type, optional Agent ID, optional Message | `hooks/aidlc-log-subagent.ts` (SubagentStop) |
### Reviewer Enforcement Events (2 events - hook-emitted)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `REVIEWER_SCOPE_BLOCKED` | A per-unit reviewer's tool call was refused for reaching into sibling units' `construction/` paths (the §12a read-scope bound) | Timestamp, Tool, Target, Stage, Unit | `hooks/aidlc-reviewer-scope.ts` (PreToolUse) |
| `REVIEW_FREEZE_BLOCKED` | A file-tool or shell `produces[]` write was refused because it would invalidate a fresh READY review receipt before the gate (the §12a terminal-receipt ordering) | Timestamp, Tool, Target, Stage, optional Unit | `hooks/aidlc-review-freeze.ts` (PreToolUse) |
### Plan Approval Enforcement Events (2 events - hook-emitted)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `PLAN_APPROVAL_BLOCKED` | A code-generation developer-agent dispatch or workspace mutation was refused because the active unit or zero-Unit stage target lacked a current, explicitly approved plan contract (stage Steps 2-3 must precede Step 4) | Timestamp, Tool, Target, Stage, Unit | `hooks/aidlc-plan-approval-guard.ts` (PreToolUse) |
| `GUARD_DISABLED` | A tool call passed the Plan Approval guard because its deterministic off-switch (the disable environment variable documented in the hooks reference) was set while a workflow existed. One row per streak: the hook appends only when the newest row in the active shard is not already this event for the same guard | Timestamp, Guard (`plan-approval-guard`), Tool | `hooks/aidlc-plan-approval-guard.ts` (PreToolUse) |
### Documents (3 events)
The DocumentKB is a **space-level** store, so all three events land in the space-level audit shard (`spaces/<space>/intents/audit/`) even for an intent-scoped document — the intent UUID is recorded as a field rather than selecting the shard. This keeps one document's history in one place across an `associate`/`dissociate` scope change. These events are provenance, **not a backup**: deleting the whole `documentkb/` tree destroys the per-document records, so document ids, tombstones, and intent links do NOT survive it — the next `sync` re-indexes the surviving originals as brand-new rows. Only a lost `index.json` alone is recoverable (`sync` rebuilds it from the per-document `metadata.json` files).
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `DOCUMENT_INDEXED` | A customer document was indexed into the DocumentKB for the first time (from `onboard`, and from `sync`'s fresh-document branch) | Timestamp, Space, Document, Source, Digest, optional Intent | `tools/aidlc-knowledge.ts` |
| `DOCUMENT_UPDATED` | An indexed document's record changed — a new revision, a re-extraction, a move, a summary, or an intent association change (from `associate`, `dissociate`, `rebind`, `summarize`, `onboard`'s edited-row branch, `sync`'s moved/changed/retried branches, and the idempotent audit-repair pass) | Timestamp, Space, Document, Change, optional Intent | `tools/aidlc-knowledge.ts` |
| `DOCUMENT_REMOVED` | The original is gone; the row became a metadata-only tombstone and extracted content was deleted (from `sync`) | Timestamp, Space, Document, Last Path, Last Digest | `tools/aidlc-knowledge.ts` |
All three are written to the **space-level** shard (`intents/audit/`), not an intent's,
even when the document carries `related_intent_ids`. A document outlives any single
intent and `associate`/`dissociate` can move its scope, so filing its provenance under
the active intent would split one document's history across shards. An
`associate`/`dissociate` that changes nothing emits **no** event: a per-call event
would fill the ledger with non-changes and break reconstruction-from-the-ledger.
### Utility Events (1 event)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `HEALTH_CHECKED` | `--doctor` completed | Timestamp, Request, Details | `tools/aidlc-doctor.ts` via the shared utility collector |
### Error/Recovery Events (2 events)
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `ERROR_LOGGED` | Tool CLI exited non-zero via `error()` | Timestamp, Tool, Command, Error | `tools/aidlc-lib.ts emitError` (called by every tool's `error()` helper) |
| `RECOVERY_COMPLETED` | User answered the compaction-awareness prompt | Timestamp, Choice, Current Stage | `tools/aidlc-state.ts acknowledge-compaction` |
### Construction Bolt Events (4 events)
Emitted only during Phase 3 (Construction). See `stage-protocol.md` Terminology for Bolt as a sprint-like iteration whose intended Unit grouping is recorded during Delivery Planning (2.9), while runtime batching comes from the Unit dependency graph. `BOLT_STARTED` / `BOLT_COMPLETED` (and swarm-path `BOLT_FAILED`) are emitted from `aidlc-bolt.ts` on the swarm / worktree path; a default gated run does not record them. `AUTONOMY_MODE_SET` is emitted by `aidlc-bolt.ts set-autonomy` after the ladder on any walk, including a default gated run.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `BOLT_STARTED` | Swarm / worktree path only: `aidlc-bolt.ts start` for one Unit and its worktree. Not emitted on a default gated stage-major run. | Timestamp, Bolt names, Batch number, Walking skeleton (true/false), optional Bolt slug, Base commit, Base Source Listing (when `--worktree`; the listing is the content-addressed raw-aware source baseline propagated from worktree creation), and Attempt Generation (team Unit claim) | `tools/aidlc-bolt.ts start` |
| `BOLT_COMPLETED` | Swarm / worktree path only: `aidlc-bolt.ts complete` for that same Unit/worktree. Does **not** close the batch — `SWARM_COMPLETED` does. | Timestamp, Bolt names, Batch number, optional Bolt slug (when --merge), optional Attempt Generation (team Unit claim) | `tools/aidlc-bolt.ts complete` |
| `BOLT_FAILED` | A Bolt failed during code-generation, or was explicitly aborted by the user | Timestamp, Failed Bolt, Error summary, optional Bolt slug (halt-and-ask correlation surface read by `aidlc-worktree info --slug`), optional Reason (`aborted` for explicit abort), optional Succeeded siblings, optional Attempt Generation (team Unit claim) | `tools/aidlc-bolt.ts fail` and `tools/aidlc-bolt.ts abort` |
| `AUTONOMY_MODE_SET` | User answered the ladder prompt after the walking skeleton | Timestamp, Mode (`autonomous` or `gated`) | `tools/aidlc-bolt.ts set-autonomy` |
### Worktree (7 events)
Emitted during Phase 3 (Construction) when Bolts run inside per-Bolt git worktrees. Worktree primitive emits `WORKTREE_*`; state fork/merge subcommands emit `STATE_*`; audit fork/merge subcommands emit `AUDIT_*`.
`Worktree path` values are project-relative (`.aidlc/worktrees/bolt-<slug>`) in
new rows. Readers resolve them against the project root and remain compatible
with legacy absolute values.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `WORKTREE_CREATED` | Per-Bolt git worktree created from main on Bolt start | Timestamp, Bolt slug, project-relative Worktree path, Branch name, Base branch, Base commit, Base Source Listing (`sha256:<hash>` over the raw-aware source listing computed from the immutable base before the audit-first create), Repo (recorded selector or `-` for the workspace root), optional Intent record and Swarm Unit/Batch/Stage/Run floor provenance | `tools/aidlc-worktree.ts` (`create`) |
| `WORKTREE_MERGED` | Bolt's worktree merged back to main on gate approval | Timestamp, Bolt slug, Worktree path, Target branch, Strategy | `tools/aidlc-worktree.ts` (`merge`) |
| `WORKTREE_DISCARDED` | Aborted Bolt's worktree explicitly removed | Timestamp, Bolt slug, Worktree path, Reason | `tools/aidlc-worktree.ts` (`discard`) |
| `STATE_FORKED` | State file forked to worktree on Bolt start | Timestamp, Bolt slug, Worktree path, Source state hash, Target state hash, optional Attempt Generation (team Unit claim) | `tools/aidlc-state.ts` (`fork`) |
| `STATE_MERGED` | Worktree's state merged back to main state on gate approval | Timestamp, Bolt slug, Worktree path, Source state hash, Target state hash, Conflict resolution | `tools/aidlc-state.ts` (`merge`) |
| `AUDIT_FORKED` | Audit log forked to worktree on Bolt start (audit-of-intent — emit precedes the byte-copy) | Timestamp, Bolt slug, Source Audit Hash, Fork Boundary, optional Attempt Generation (team Unit claim) | `tools/aidlc-audit.ts` (`audit-fork`) |
| `AUDIT_MERGED` | Worktree's audit entries appended to main audit on gate approval; per-Bolt entry order preserved, cross-Bolt order reflects merge-completion order. The review records named by the delta's paired `REVIEW_COMPLETED` rows are carried into the main intent record first, digest-verified, or the merge refuses | Timestamp, Bolt slug, Entries Merged, Source Audit Hash, Fork Boundary, Fork Timestamp | `tools/aidlc-audit.ts` (`audit-merge`) |
### Practices (4 events)
Emitted by the Inception stage `practices-discovery` and by the Construction orchestrator at runtime. The stage emits at the affirmation gate; the orchestrator emits at runtime via `--type empty` (fallback advisory) and `--type override` (discriminator-field for the bolt-plan-marker-conflict path).
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `PRACTICES_DISCOVERED` | Greenfield or brownfield lead draft, three support contributions, human interview, and lead integration completed; drafts await affirmation | Timestamp, sources scanned, drafts produced | `tools/aidlc-state.ts` `practices-event --type discovered` |
| `PRACTICES_AFFIRMED` | Team approved practices at the practices-discovery affirmation gate; content promoted to `aidlc/spaces/<active-space>/memory/team.md` and `project.md` | Timestamp, affirming user, sections written, mandated/forbidden rules appended | `tools/aidlc-state.ts` `practices-promote` |
| `PRACTICES_OVERRIDE` | Cross-row promotion failed during practices-discovery affirmation, OR walking-skeleton stance from active-space `team.md` overrode bolt-plan's marker for the current Bolt | Timestamp, Reason (discriminator); per-path field set: write-failure path emits Reason + Failure detail only (no Bolt fields); bolt-plan-marker-conflict path emits Reason + Bolt slug + Practices Stance + Bolt-Plan Marker. The two field sets do not overlap, so doctor filters by `Reason` and routes by either name family — `write-failure-*` for the affirmation promotion path, `bolt-plan-marker-conflict` for the orchestrator runtime path | `tools/aidlc-state.ts` `practices-promote` (write-failure path); `tools/aidlc-state.ts` `practices-event --type override` (bolt-plan-marker-conflict path — discriminator-field disambiguation, no separate event) |
| `PRACTICES_SECTION_EMPTY` | Orchestrator read a practices section that returned empty; falling back to org defaults (advisory-only) | Timestamp, Section name, Fallback source | `tools/aidlc-state.ts` `practices-event --type empty` |
### Merge Dispatch (3 events)
Emitted when Construction's Bolt-merge step calls aidlc-pipeline-deploy-agent via Task to determine the merge strategy from team practices prose. Emitted via the `aidlc-bolt dispatch-event` subcommand. The orchestrator brackets each aidlc-pipeline-deploy-agent dispatch — pre-call INVOKED, post-call RETURNED on successful parse, FALLBACK on timeout/malformed-YAML. Audit-of-intent semantic: INVOKED emits before the LLM Task call (no disk side-effect for the dispatch itself; reconciliation by slug + timestamp window). Doctor reconciles orphan INVOKED rows.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `MERGE_DISPATCH_INVOKED` | Orchestrator dispatched aidlc-pipeline-deploy-agent with current practices section + Bolt context | Timestamp, Bolt slug, Practices section excerpt; optional Pinned OID + Attempt Generation + Pin Transaction for Unit merge gates | `tools/aidlc-bolt.ts` `dispatch-event --event MERGE_DISPATCH_INVOKED` |
| `MERGE_DISPATCH_RETURNED` | Agent returned parsed YAML with strategy, target branch, confidence, notes | Timestamp, Bolt slug, Strategy, Target branch, Confidence, Notes; optional Pinned OID + Attempt Generation + Pin Transaction | `tools/aidlc-bolt.ts` `dispatch-event --event MERGE_DISPATCH_RETURNED` |
| `MERGE_DISPATCH_FALLBACK` | Agent timed out or returned malformed YAML; orchestrator fell back to org defaults — critical observability hook | Timestamp, Bolt slug, Fallback reason, Defaults applied; optional Pinned OID + Attempt Generation + Pin Transaction | `tools/aidlc-bolt.ts` `dispatch-event --event MERGE_DISPATCH_FALLBACK` |
### Sensor Events (5 events)
Emitted by the deterministic-sensor system. The sensor dispatcher emits the four `SENSOR_*` events; the paired-coverage doctor row emits `GUARDRAIL_LOADED` with `Scope: all`, because doctor reads the full resolved guardrail set without an active stage (the per-workflow org → project → phase → stage scoping in the When-clause below describes the steady-state loader, not doctor's unscoped read). `fire_on: write` sensors run from the PostToolUse hook on matching writes; `fire_on: gate` sensors run once per existing declared deliverable before `gate-start` opens the first gate, `revise` re-enters it, or the approve-time revision backstop attempts recovered re-entry. Blocking severity is enforced only for gate-fired sensors in this release; write-fired failures remain advisory.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `SENSOR_FIRED` | Dispatcher invoked a sensor against a stage output, from either a matching PostToolUse Write/Edit or gate-boundary deliverable dispatch | Timestamp, Fire id, Sensor ID, Stage slug, Output path | `tools/aidlc-sensor.ts` `fire` |
| `SENSOR_PASSED` | Sensor completed and reported no findings (also: tool-unavailable, script-error fall-through — see Note footnote) | Timestamp, Fire id, Sensor ID, Stage slug, Output path, Duration ms | `tools/aidlc-sensor.ts` `fire` |
| `SENSOR_FAILED` | Sensor completed and reported findings; detail file written at `<record>/.aidlc-sensors/<stage-slug>/<sensor-id>-<fire-id>.md` | Timestamp, Fire id, Sensor ID, Stage slug, Output path, Detail path, Findings count | `tools/aidlc-sensor.ts` `fire` |
| `SENSOR_BUDGET_OVERRIDE` | Sensor exceeded its configured cap (registry / binding / depth-derived per the three-layer cap model) and was terminated or skipped | Timestamp, Fire id, Sensor ID, Stage slug, Output path, Cap layer, Cap value, Observed value | `tools/aidlc-sensor.ts` `fire` |
| `GUARDRAIL_LOADED` | Guardrail loader resolved the scope-hierarchical guardrail set for the active workflow (org → project → phase → stage); doctor's paired-coverage check reads from this event | Timestamp, Scope, Path, Rule count | `tools/aidlc-utility.ts` |
> The `Note` field on `SENSOR_PASSED` is optional. It carries `tool-unavailable` when the per-sensor script's underlying binary isn't on PATH, or `script-error: <reason>` for spawn-failure / non-zero exit / malformed JSON / detail-write failure paths. Those remain advisory for write-fired sensors, but a blocking gate binding requires a verified pass, so any `Note`, dispatcher failure, malformed verdict, or budget override refuses gate entry. Pair correlation is via `Fire id` (echoed verbatim from `SENSOR_FIRED` to the terminal row); `Output path` alone does not disambiguate repeated write dispatches or a later gate dispatch of the same sensor + stage + path tuple.
> **Pair by `Fire id`, not by audit-row index.** Write dispatch can fan out one tool call across several sensors, while gate dispatch fans each gate-bound sensor across every existing deliverable. Terminal rows may interleave by spawn duration, so `findAllEvents("SENSOR_FIRED")[i]` does NOT pair with `findAllEvents("SENSOR_PASSED")[i]` by index. Audit-walking consumers (the `sensor_firings[]` populator, doctor, designer) MUST match terminal rows to FIRED rows via the 8-hex `Fire id` correlator.
### Learning Loop (3 events)
Emitted by stage-protocol §13 (Learnings Ritual). The runtime-graph compile emits `MEMORY_EMPTY` when a just-approved stage's memory.md has zero non-blank entries under the four standard headings. The learning-gate tool emits `RULE_LEARNED` when the user keeps a surfaced or free-text learning (a learning IS a practice — it lands as a practice line under the routed heading in `{project,team}.md`) and `SENSOR_PROPOSED` when a learning installs a sensor binding (manifest + originating stage `sensors:` frontmatter). Doctor reads `MEMORY_EMPTY` rows over time to detect systematic diary-skipping across stages.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `MEMORY_EMPTY` | A stage approval triggered a runtime-graph compile and the stage's memory.md had zero non-blank entries under any of the four §13 headings | Timestamp, Stage | `tools/aidlc-runtime.ts compile` |
| `RULE_LEARNED` | The learning gate persisted a kept learning as a practice line under the routed heading in `{project,team}.md` | Timestamp, Stage, Candidate-ID, Content-Hash, Destination, Heading, Source | `tools/aidlc-learnings.ts persist` |
| `SENSOR_PROPOSED` | The learning gate scaffolded a project-tier sensor manifest and bound it to the originating stage's `sensors:` frontmatter | Timestamp, Stage, Candidate-ID, Sensor ID, Manifest path, Matches, Destinations, Source | `tools/aidlc-learnings.ts persist` |
### Swarm (7 events)
Six events emit from the swarm referee `aidlc-swarm.ts` — the deterministic verdict surface the conductor consults — and `SWARM_SOURCE_MERGED` emits from `aidlc-worktree.ts` after application source lands in main. The referee is stateless (no iteration counter): `prepare` captures the exact stage-attempt token, stamps the Unit/Batch/Stage/Run floor into `WORKTREE_CREATED` plus worktree metadata, forks the per-unit worktrees, and emits `SWARM_STARTED` with both the prepared batch and the full attempt-bound Unit obligation set (and `SWARM_DEGRADED` when the conductor reports a loud downgrade); `finalize` requires that prepared token to still match the current attempt before merging, then preserves it on each convergence row. Completion compares the live authored DAG with that durable obligation set and refuses shrink or expansion until the DAG is restored or the stage attempt restarts. A worktree prepared by the immediately preceding unstamped release remains finalizable only when its frozen audit prefix proves the matching legacy `SWARM_STARTED`, `BOLT_STARTED`, `STATE_FORKED`, and `AUDIT_FORKED` sequence, the fork hash still matches, all three worktree mirrors exist, and the derived frozen attempt still equals the current attempt; this compatibility path never emits a backdated start or adopts a stale worktree. It re-verifies the conductor's claimed-converged set, snapshots the exact declared record artifacts plus the bound `source-manifest.json`, serialised-merges those records and AIDLC metadata for the genuine passes, and emits the per-Unit pair (`SWARM_UNIT_CONVERGED` / `SWARM_UNIT_FAILED` — except a converged unit whose record/metadata merge-back failed, which gets neither row until a finalize retry merges it), the per-failed-Unit baton row (`SWARM_BATON_RETURNED`), and the batch tally (`SWARM_COMPLETED`). Source merge then requires authority correlated to durable worktree provenance and the current Bolt/batch/stage attempt, links the main checkout from the stage-entry baseline or the prior attempt's validated final aggregate recorded on `GATE_REJECTED`, and emits `SWARM_SOURCE_MERGED`. Rejection never replaces the completion baseline. The opening link is revalidated against that trusted predecessor; later links must form an exact fingerprint chain. A pre-binding worktree whose convergence predates immutable source fields retains the historical branch merge; modern incomplete authority fails closed and does not advance batch routing. If Git mutation lands but `SWARM_SOURCE_MERGED` cannot be appended, the command preserves the worktree and reports a non-retryable `[merge-succeeded:<sha>]` failure: do not rerun the merge, because no authenticated recovery receipt exists; restart the stage attempt or use the explicit source-freshness off-switch only with human approval. If `SWARM_SOURCE_MERGED` did land and only subsequent worktree, branch, or retained-ref cleanup failed, rerunning the same merge performs cleanup-only reconciliation without reapplying source or duplicating authority. The `check` subcommand emits nothing — it is an advisory verdict that informs the conductor's retry decision. Because the loop and its cap live in the driver, the per-Unit rows carry no `Iterations` / `Cap value` fields.
| Event | When | Required Fields | Emitter |
|-------|------|-----------------|---------|
| `SWARM_STARTED` | Swarm referee `prepare` captured the exact attempt and forked a batch of dependency-linked Units | Timestamp, Batch number, Unit names, Unit obligations, Concurrency cap, Stage, Run floor | `tools/aidlc-swarm.ts` |
| `SWARM_UNIT_CONVERGED` | A swarm Unit re-verified green (and untampered), its exact declared record artifacts plus bound source manifest landed, and its AIDLC metadata merge-back landed; unless explicitly bypassed, finalize also verified the current global/unit source bindings and raw-aware base-to-worktree footprint against reviewed source-manifest claims | Timestamp, Batch number, Unit name, Stage, Run floor, optional Source Fingerprint and Source Commit (immutable reviewed source accepted by `aidlc-worktree merge`), or `Source Freshness Bypass: true` when finalize explicitly used `AIDLC_SKIP_SOURCE_FRESHNESS=1` (the freshness/unit/footprint guarantees do not apply and merge must repeat the switch) | `tools/aidlc-swarm.ts` |
| `SWARM_SOURCE_MERGED` | The immutable reviewed source for one current-attempt swarm Unit landed in main and extended the aggregate source chain | Timestamp, Batch number, Unit name, Stage, Run floor, Previous Source Fingerprint, Source Fingerprint, Source Commit, Merge commit, Repo (recorded selector or `-` for the workspace root) | `tools/aidlc-worktree.ts merge` |
| `SWARM_UNIT_FAILED` | A swarm Unit failed the `finalize` re-verify (not claimed, claimed-but-red, or tampered) | Timestamp, Batch number, Unit name, Reason | `tools/aidlc-swarm.ts` |
<!-- Reason for a CLAIMED-but-red / tampered unit is always the tool's own verdict (`error`); for a DECLINED (unclaimed) unit it is the conductor's typed attribution via `finalize --reasons` (`unsatisfiable` / `budget-exhausted` / `cap-exhausted`, defaulting to `cap-exhausted`) — the tool records the conductor's knowledge call, it does not judge unsatisfiability itself (D-I). -->
| `SWARM_BATON_RETURNED` | A swarm Unit returned the baton to the conductor for orchestrator-mediated coordination | Timestamp, Batch number, Unit name, Reason | `tools/aidlc-swarm.ts` |
| `SWARM_COMPLETED` | All Units in the batch finished (converged or failed); batch closed | Timestamp, Batch number, Converged count, Failed count | `tools/aidlc-swarm.ts` |
| `SWARM_DEGRADED` | `AIDLC_USE_SWARM=1` was requested but the Workflow tool was unavailable, so the conductor ran the subagent floor (loud-degrade) | Timestamp, Batch number, Requested driver, Fallback driver | `tools/aidlc-swarm.ts` |
## Hook-Generated Format
Hooks emit events through the same library emitter as orchestrator-driven emissions (`appendAuditEntry` from `tools/aidlc-audit.ts`). Hook-emitted events are first-class taxonomy members (`ARTIFACT_CREATED`, `ARTIFACT_UPDATED`, `SUBAGENT_COMPLETED`, all `SESSION_*`) — there is no longer a separate "free-form hook entry" format. A hook with no active workflow in `cwd` is a no-op; session events only append to a workflow's audit.md when one exists.
The public `aidlc-audit.ts append` CLI is a diagnostic escape hatch, not the canonical emit path: it refuses authority-bearing receipts (`STAGE_COMPLETED`, `HUMAN_TURN`, `GATE_APPROVED`, `GATE_REJECTED`, `QUESTION_ANSWERED`, `REVIEW_REQUESTED`, `REVIEW_COMPLETED`, `PIPELINE_LINK_COMPLETED`, `ARTIFACT_REUSED`, `SWARM_STARTED`, `SWARM_UNIT_CONVERGED`, `SWARM_SOURCE_MERGED`, `AUTONOMY_MODE_SET`, `UNIT_OWNERSHIP_SET`, `UNIT_GATE_RHYTHM_SET`, `UNIT_STARTED`, `UNIT_PAUSED`, `UNIT_RESUMED`, `UNIT_COMPLETED`, `UNIT_MERGED`, `DOCUMENT_INDEXED`, `DOCUMENT_UPDATED`, `DOCUMENT_REMOVED`), which only their owning tool or hook may emit. Field names must be printable single-line labels matching the audit field grammar; values have every line terminator escaped. `append-raw` likewise refuses a body carrying an `**Event**:` line naming a taxonomy event and refuses line-breaking headings.
## Format Standards
- All timestamps: ISO 8601 format (YYYY-MM-DDTHH:MM:SSZ)
- Generate fresh timestamp for EACH entry via `date -u +"%Y-%m-%dT%H:%M:%SZ"` (tools do this automatically)
- Append-only — NEVER modify or delete existing entries
- No sensitive data (credentials, PII, secrets)
- Human decisions recorded verbatim — NEVER summarize
## Entry Format
### Standard Format
```
## [Event Heading]
**Timestamp**: [ISO timestamp]
**Event**: [Event type from table above]
**Stage**: [Stage slug — optional, context-dependent]
**Details**: [Event-specific content]
---
```
### Error Format
```
## Error: [Brief Description]
**Timestamp**: [ISO timestamp]
**Severity**: [Critical/High/Medium/Low]
**Type**: [Parse error/Missing artifact/State corruption/Validation failure]
**Description**: [What went wrong]
**Resolution**: [Action taken]
---
```
### Recovery Format
```
## Recovery: [Brief Description]
**Timestamp**: [ISO timestamp]
**Issue**: [What triggered recovery]
**Steps**: [Numbered recovery actions]
**Outcome**: [Successful/Partial/Failed]
---
```
## Validation basis on `STAGE_COMPLETED`
Main-workflow `STAGE_COMPLETED` entries may include a compact schema-3 receipt:
```text
**Validation Basis**: {"schema":3,"graphContract":"sha256:...","projectType":"brownfield","inputs":[{"artifact":"code-summary","producer":"code-generation","required":true,"instanceCount":12,"presentCount":12,"structureHash":"sha256:...","contentHash":"sha256:..."}],"outputs":[...]}
```
The resolver first computes the concrete runtime artifact-instance set using
the Bolt DAG, unit kinds, `produces_kinds`, and canonical filename aliases. The
receipt then stores deterministic stage-level aggregate fingerprints rather
than every Unit/path row. Optional inputs are included only when at least one
instance was present when completion was reported. This is a report-time
snapshot: it does not prove which bytes the stage actually read. "Observed"
dependency means recorded in this receipt, not runtime read instrumentation;
changes before capture become the baseline and changes after capture are
detectable.
If receipt capture is unavailable, the owning completion tool still appends a
receipt-less `STAGE_COMPLETED` with `Validation Warning`. That completion is
untracked and advisory rather than a reason to abort the state transition.
The latest receipt remains useful as dependency evidence after a stage is
reopened, while only a completion in the current attempt counts as current
tracking for the stage itself. Earlier schemas and receipt-less completions fail
open until the stage completes with this schema. Schema 2 is deliberately
untracked because it cannot distinguish its old zero-instance resolution from
the stage-level zero-Unit resolution introduced with schema 3.
@@ -0,0 +1,30 @@
# Brownfield Safeguards
For any stage that modifies existing code or infrastructure, these safeguards apply.
## Safeguard Matrix
| Safeguard | What It Does | When Applied |
|-----------|-------------|--------------|
| **Blast Radius Analysis** | Identifies affected files/components and their downstream dependents | Before code generation (stage `code-generation`, 3.5) |
| **Diff Preview** | Shows exact proposed changes before applying | Before any file modification |
| **Test Baseline** | Runs existing tests BEFORE changes to establish baseline | Before code generation (stage `code-generation`, 3.5) |
| **Test Validation** | Runs existing tests AFTER changes to confirm nothing broke | After code generation (stage `build-and-test`, 3.6) |
| **Impact Analysis** | Documents affected APIs, components, and dependencies | During reverse engineering (stage `reverse-engineering`, 2.1) and code generation (3.5) |
| **Rollback Plan** | Documents how to undo changes if needed | Before deployment (stage `deployment-execution`, 4.3) |
## Blast Radius Analysis Template
Before modifying existing code:
1. List all files that will be changed
2. For each file, identify: imports/consumers, test files, configuration references
3. Classify impact: low (isolated change), medium (affects 2-3 dependents), high (cross-cutting)
4. Present impact summary to user before proceeding
## Test Baseline Protocol
1. Run full test suite before ANY code changes
2. Record: total tests, passing, failing, skipped, coverage %
3. After code generation, re-run full test suite
4. Compare: new failures = regressions introduced by changes
5. If regressions found, fix before proceeding
@@ -0,0 +1,34 @@
# Team Knowledge
Add markdown files here to extend AI-DLC agent behavior for your project.
## How It Works
Files placed in these directories are loaded by agents during every stage (after the built-in methodology). Use them for:
- Company coding standards and conventions
- Architectural decisions and tech stack preferences
- Domain-specific terminology and business rules
- Project-specific patterns and anti-patterns
## Directory Structure
> This table is a snapshot — the authoritative mapping lives in each agent's frontmatter at `.aidlc/agents/*.md`.
| Directory | Purpose | Example files |
|-----------|---------|---------------|
| `shared/` | Team-wide standards | coding-standards.md, api-conventions.md |
| `aidlc-architect-agent/` | Architecture decisions | tech-stack.md, infrastructure-preferences.md |
| `aidlc-developer-agent/` | Coding patterns | db-conventions.md, error-handling.md |
| `aidlc-quality-agent/` | Testing standards | test-strategy.md, coverage-requirements.md |
| `aidlc-design-agent/` | UX/UI guidelines | design-system.md, accessibility.md |
| `aidlc-product-agent/` | Product context | roadmap.md, personas.md |
| `aidlc-devsecops-agent/` | Security policies | security-baseline.md, compliance-rules.md |
| `aidlc-operations-agent/` | Ops runbooks | monitoring.md, incident-response.md |
| `aidlc-compliance-agent/` | Compliance standards | data-governance.md, audit-requirements.md |
| `aidlc-aws-platform-agent/` | Cloud infrastructure | account-structure.md, service-limits.md |
| `aidlc-pipeline-deploy-agent/` | CI/CD configuration | pipeline-standards.md, deployment-gates.md |
| `aidlc-delivery-agent/` | Project management | sprint-cadence.md, definition-of-done.md |
## Format
Any `.md` file placed in a directory is loaded. No special naming required. Keep files focused — one topic per file works best.
@@ -0,0 +1,14 @@
<!-- INVARIANT: examples are single-line HTML comments so a fresh template parses to total=0 (MEMORY_EMPTY). Do NOT un-comment or split across lines. t100 guards this. -->
> This file is kept up to date automatically while the stage runs. Add observations at the review step, not by editing here directly.
## Interpretations
<!-- example: 2026-05-29T10:14:32Z — chose REST over GraphQL; the consuming team only needs CRUD, revisit if subscriptions land -->
## Deviations
<!-- example: 2026-05-29T10:14:32Z — skipped the optional caching layer the stage prose suggested; the dataset is small enough that it adds risk -->
## Tradeoffs
<!-- example: 2026-05-29T10:14:32Z — picked TDD over BDD this run; the team is unit-first and the domain is well-understood -->
## Open questions
<!-- example: 2026-05-29T10:14:32Z — confirm the retention window with compliance before the next stage hardens the schema -->
@@ -0,0 +1,175 @@
# Reading Active-Space Rule Files
> **Audience**: any agent that needs to read team-affirmed practices from
> `aidlc/spaces/<active-space>/memory/`.
> **Owner of this file**: framework. Cited by
> `aidlc-pipeline-deploy-agent/branching-strategies.md` and by other agents
> that adopt practices-aware behaviour.
The rules namespace resolves through a strict-additive five-layer chain
at workflow start: `org → team → project → phase → stage`. The first three
files are `aidlc/spaces/<active-space>/memory/{org,team,project}.md`;
`phases/<phase>.md` attaches because the stage's frontmatter `phase: <name>`
selects it. Stage rules are reserved for future use. The compile bakes the
resolved chain into each stage node's `rules_in_context`; every applicable
rule appears in the chain - nothing drops at runtime. At stage entry the
engine delivers each substantive rule file through an ordered sequence of
bounded `load-steering` directives before `run-stage` (empty templates are
dropped per § 1; there is no size-based path fallback). This file documents
how to read those layers safely when an agent needs a practice outside that
delivery sequence (tool shaping, question prompts).
---
## 1. Empty-template detection
A rule section is **empty** when every non-blank line in its body begins
with `<!--` or is whitespace. The template ships with HTML-comment
placeholders illustrating affirmed-state examples; until the team has
affirmed practices, every section is empty by this rule.
When you read a section and find it empty, fall back to the next layer
(see § 3). Do not parse the example prose — the comments exist for human
readers, not for agent inference.
Example empty section:
```
## Way of Working
<!-- Populated by practices-discovery affirmation. Example after affirmation:
"We squash-merge to `main`. Trunk-based; feature branches resolve in 1-2 days." -->
```
Example populated section:
```
## Way of Working
We squash-merge to `main`. Trunk-based; feature branches resolve in 1-2 days.
```
The first body line being `<!--` is the empty signal; the second body line
being prose is the populated signal.
---
## 2. Semantic-topic matching
Heading shapes can drift between `team.md`, `org.md`, `project.md`, and the
per-agent KB that consumes them. Match by **topic**, not by exact-string
heading.
For the **way of working / branching / merge** topic:
- Prefer an exact `## Way of Working` heading.
- Fall back to any `## ` heading containing `branch`, `merge`, or `way`
(case-insensitive).
For the **walking skeleton** topic:
- Prefer `## Walking Skeleton`.
- Fall back to any heading containing `skeleton` (case-insensitive).
For the **testing** topic:
- Prefer `## Testing Posture`.
- Fall back to any heading containing `test` (case-insensitive).
- For Code Generation, resolve the explicit `Methodology` and `Ordering`
statements independently from coverage, tooling, integration, or scope
notes. A narrower note that says nothing about cadence remains additive and
does not erase a broader affirmed methodology. A contradictory narrower
methodology is an error under the strict-additive model.
For the **deployment** topic:
- Prefer `## Deployment`.
- Fall back to any heading containing `deploy` or `release`
(case-insensitive).
For the **code style** topic:
- Prefer `## Code Style`.
- Fall back to any heading containing `style` or `format`
(case-insensitive).
When multiple candidate headings match by fallback, prefer the first
occurrence in document order.
---
## 3. Fallback chain
For a decision that needs one operative practice statement, inspect the
active-space descriptive layers from narrowest to broadest and return the
first non-empty, non-conflicting section:
1. **`aidlc/spaces/<active-space>/memory/project.md`** — project-specific
specialisation.
2. **`aidlc/spaces/<active-space>/memory/team.md`** — team-affirmed practices.
3. **`aidlc/spaces/<active-space>/memory/org.md`** — framework defaults written
in team voice.
4. **Hardcoded defaults** — used only when all three layers are empty
(greenfield first run before practices-discovery has run, or when the
project SKIPs practices-discovery).
This topic selection does not erase broader rules: the runtime still loads all
applicable layers. A narrower statement that contradicts broader policy is an
admission error, not an override.
Hardcoded defaults are:
| Topic | Default |
|---|---|
| Way of Working | trunk-based development; base `main`, target `main`; squash-merge |
| Walking Skeleton | scope-dependent; the active scope file's `skeleton:` field supplies the default ceremony stance |
| Testing Posture | test-after ordering when no methodology is affirmed; the test-strategy axis governs volume |
| Deployment | trunk-based with on-merge staging deploy; production gate is human-approved |
| Code Style | defer to project linter/formatter configuration |
When the fallback chain has to descend to layer 4, emit
`PRACTICES_SECTION_EMPTY` (advisory-only) so doctor and downstream
observability can flag projects running on framework defaults vs affirmed
team practices.
---
## 4. Read protocol (pseudocode)
```
def read_practice(topic):
for layer in [
aidlc/spaces/<active-space>/memory/project.md,
aidlc/spaces/<active-space>/memory/team.md,
aidlc/spaces/<active-space>/memory/org.md,
]:
section = match_section(layer, topic) # § 2
if section and not is_empty(section): # § 1
return section
return hardcoded_default(topic) # § 3
```
The protocol is intentionally synchronous and side-effect-free. Agents
call this when shaping a tool invocation; the orchestrator calls it when
shaping a question prompt. Both surfaces use the same fallback chain so
user-visible behaviour stays consistent across the dispatch.
---
## 5. Worked example — aidlc-pipeline-deploy-agent reading way-of-working
The orchestrator dispatches `aidlc-pipeline-deploy-agent` at Bolt-create time.
The agent's job is to map team intent to `aidlc-worktree create --slug
<slug> --base <branch>`. It reads:
1. `project.md` `## Way of Working` → empty (fresh template).
2. `team.md` `## Way of Working` → empty (fresh template).
3. `org.md` `## Way of Working` → "trunk-based; base `main`, target
`main`; squash-merge".
4. Returns `{base: "main", strategy: "squash"}` to the orchestrator
alongside the agent's invocation of `aidlc-worktree`.
If `team.md` `## Way of Working` had read "We use GitFlow with
`develop` as the integration branch", the agent would map that to
`--base develop` instead — same fallback chain, populated layer wins.
If all three layers were empty (a project that SKIPped practices-discovery),
the agent would emit `PRACTICES_SECTION_EMPTY` and apply the hardcoded
default — `{base: "main", strategy: "squash"}`. Doctor surfaces this so
the team can run practices-discovery later if they want their own
practices captured.
@@ -0,0 +1,91 @@
# AI-DLC State Tracking
This document defines the `aidlc-state.md` section and field contract. The
engine writes the concrete state file and enumerates stages from the compiled
stage graph plus scope grid; this template must not hand-list shipped stages.
The exact initial description is JSON-encoded as one string beside the state
file in `<record>/project-description.json`; the `Project` field below is its
safe single-line preview.
Authoritative generated views:
- Stage graph: `aidlc engine gen stage-table`
- Scope grid: `aidlc engine gen scope-table`
## Project Information
- **Project**: [single-line project description preview]
- **Project Description Source**: project-description.json
- **Project Type**: [Greenfield/Brownfield]
- **Scope**: [scope slug from compiled scope grid]
- **Start Date**: [ISO 8601 timestamp]
- **State Version**: 8
- **Active Agent**: [current lead agent slug]
- **Worktree Path**: [empty when not in a worktree]
- **Bolt Refs**: [empty list or comma-separated bolt slugs]
- **Practices Affirmed Timestamp**: [ISO 8601 timestamp on affirmation]
## Scope Configuration
- **Stages to Execute**: [comma-separated stage numbers included in scope]
- **Stages to Skip**: [comma-separated stage numbers with reasons, or none]
- **Depth**: [Minimal/Standard/Comprehensive]
- **Test Strategy**: [Minimal/Standard/Comprehensive]
- **Change Control**: [strict/relaxed, then its source in parentheses: `(from scope <name>)`, `(from <layer>.md)`, or `(set by you)`; written at intent creation with the resolved value, rewritten by `/aidlc --change-control` or the plain-chat request, read by value only]
## Workspace State
- **Project Root**: [project-relative path, normally `.`; re-derived at runtime, never trusted as an absolute path]
- **Languages**: [detected languages]
- **Frameworks**: [detected frameworks]
- **Build System**: [detected build system]
## Execution Plan Summary
- **Total Stages**: [count of EXECUTE stages]
- **Completed**: [count of completed EXECUTE stages]
- **In Progress**: [current stage slug]
## Runtime State
- **Revision Count**: [integer]
- **Unit Ownership**: [solo/team; optional, exact `team` activates the derived grid]
- **Unit Gate Rhythm**: [per-stage/unit-end; optional, defaults to per-stage under team ownership]
## Phase Progress
<!-- Status values: Pending, Active, Verified, Skipped -->
- **[Phase]**: [Pending/Active/Verified/Skipped]
## Stage Progress
<!-- Checkbox states: [ ] pending, [-] in-progress, [?] awaiting approval, [R] revising, [x] completed, [S] skipped -->
The engine emits one phase heading per compiled phase, then one checkbox row per
compiled stage in that phase:
### [PHASE] PHASE
- [ ] stage-slug — [EXECUTE/SKIP: reason]
## Unit Progress
Present only when `Unit Ownership: team` and `Construction Iteration:
unit-major`. This table is an engine-owned, derived projection of the Unit DAG,
artifact coverage, lifecycle receipts, and unit gate events. It is rewritten on
every `next`; hand edits are never routing or completion evidence.
| unit | owner | [per-unit Construction stage columns in graph order] | gate |
| --- | --- | --- | --- |
| [Unit name] | - | [[ ]/[-]/[?]/[R]/[x]/[S] per stage] | [[ ]/[-]/[?]/[R]/[x]] |
The stage columns use the same checkbox vocabulary as `## Stage Progress`.
`owner` remains `-` until the claim increment supplies ownership. `gate`
summarizes the current per-stage gates or the unit-end gate, depending on Unit
Gate Rhythm. Stage Progress rows are derived complete only when their Unit
Progress column and required team gates are complete.
## Current Status
- **Lifecycle Phase**: [READY/INITIALIZATION/IDEATION/INCEPTION/CONSTRUCTION/OPERATION]
- **Current Stage**: [stage slug or status text]
- **Next Stage**: [next stage slug or none]
- **Status**: [Running/Completed]
- **Construction Autonomy Mode**: [unset/autonomous/gated]
- **Last Updated**: [ISO 8601 timestamp]
## Session Resume Point
- **Last Completed Stage**: [stage slug]
- **Next Action**: [what to do next]
- **Pending Artifacts**: [any incomplete artifacts or none]
@@ -0,0 +1,74 @@
# Automatic Verification — Element-Level Traceability
`<record>` below is the active workflow record:
`aidlc/spaces/<active-space>/intents/<active-intent>`, selected by
`aidlc/active-space` (default `default`) and
`aidlc/spaces/<active-space>/intents/active-intent`.
## Per-Stage Coverage
Stages that transform requirements, stories, designs, or code produce a
declared `traceability.json` artifact. The `traceability` sensor validates each
file when it is written and the stage artifact contract makes omission visible
to the normal directive, review, and completion evidence paths.
Every file uses stable IDs:
| Prefix | Meaning | Example |
|--------|---------|---------|
| `FR{n}` / `FR{n}.{m}` | Functional requirement | `FR1`, `FR1.2` |
| `NFR{n}` | Inception non-functional requirement | `NFR2` |
| `US{n}.{m}` | User story | `US1.3` |
| `AC{n}.{m}.{seq}` | Acceptance criterion | `AC1.3.2` |
| `U{n}` / `u{n}-{description}` | Unit ID / construction directory | `U1`, `u1-auth` |
| `BR{group}.{seq}` | Business rule | `BR1.1` |
| `NFRx.y` | Detailed NFR requirement | `NFR2.1` |
The JSON shape is:
```json
{
"stage": "functional-design",
"unit": "u1-auth",
"upstream_ids": ["AC1.1.1"],
"coverage": [
{ "id": "AC1.1.1", "status": "OK", "target": "BR1.1" }
],
"reverse": [
{ "id": "BR1.3", "status": "N/A", "target": "technical validation rule" }
]
}
```
Valid statuses are `OK`, `GAP`, `ORPHAN`, `Deferred`, and `N/A`. `OK`,
`Deferred`, and `N/A` require a non-empty target or justification. The sensor
cross-checks upstream IDs from source artifacts, verifies deterministic targets
where possible, and derives functional-design business-rule orphans.
## When Verification Runs
| Trigger | What's Checked |
|---------|---------------|
| **Ideation → Inception** | Intent → Scope → Intent Backlog consistency; all scope items have feasibility backing |
| **Inception → Construction** | Requirements → Stories → Architecture alignment; all stories trace to requirements; architecture covers all stories |
| **Construction → Operation** | Architecture → Code → Tests alignment; all code traces to design; test coverage against acceptance criteria |
| **On demand** | Human can request verification at any point |
| **Stage output write** | Validate the stage's element-level coverage and targets |
## Phase Check Output
Each phase boundary check produces `<record>/verification/phase-check-<phase>.md`:
- Coverage percentages (requirements with stories, stories with components, etc.)
- Warnings (incomplete mappings)
- Consistency checks (no contradictions between phases)
- Human approval checkbox
## Verification Process
1. Read every declared traceability artifact from the completed phase
2. Rebuild the traceability chain from stable IDs
3. Identify gaps, invalid targets, derived orphans, and contradictions
4. Generate the verification report
5. Present the result for review
6. Let the engine record `PHASE_VERIFIED` in the active record's
`audit/<host>-<clone-id>.md` shard; do not append it manually
@@ -0,0 +1,91 @@
# `aidlc-worktree info` — Output Schema
Pinned schema and exit-code contract for the `info` subcommand. The orchestrator's halt-and-ask prose at `SKILL.md` reads this output to interpolate the worktree path and branch name into the structured-question prompt body for code-generation-failure halt-and-ask.
This schema is the contract between the tool (deterministic) and the LLM (prose composition). Future changes to the JSON shape must update this file in the same commit.
## Usage
```
aidlc engine worktree info --slug <kebab-slug>
```
The slug is the kebab-case Bolt identifier threaded through every worktree command for that Bolt (`create`, `verify`, `merge`, `discard`). See `SKILL.md` per-Bolt loop "Slug derivation" paragraph for the `name → slug` transformation.
## Exit codes
| Exit | Meaning | stdout | stderr |
|------|---------|--------|--------|
| 0 | Hit — JSON emitted | JSON object (see below) | (empty) |
| 1 | Miss — no `WORKTREE_CREATED` for slug, OR malformed block | (empty) | one-line error message |
The exit-code contract mirrors `verify`'s semantics: non-zero is the halt signal. The orchestrator's prose treats any non-zero exit as "no worktree to render" and falls back to the carve-out failure shape (verify-failed or dev-rejection) — but in practice this is unreachable for the wired invocation path (code-generation failure at Step 1 always has `WORKTREE_CREATED` in audit by Step 0).
## JSON output shape (exit 0)
```json
{
"slug": "onboarding-wizard",
"path": "/Users/dev/project/.aidlc/worktrees/bolt-onboarding-wizard",
"branch_name": "bolt-onboarding-wizard",
"audit_timestamp": "2026-05-18T12:34:56Z",
"merge_held": false
}
```
Field semantics:
- **`slug`** — echoes the input `--slug` flag verbatim. The slug is the bare kebab-case identifier (e.g. `onboarding-wizard`); the `bolt-` prefix on `path` and `branch_name` is added by `lib.ts:139` `worktreePath()` and the `aidlc-worktree create --slug <slug>` invocation. See SKILL.md per-Bolt loop "Slug derivation" paragraph for the `name → slug` transformation that produced the bare slug. The orchestrator uses this field to confirm correlation, not to pick a different one.
- **`path`** — absolute filesystem path of the worktree hosting the Bolt at `<projectDir>/.aidlc/worktrees/bolt-<slug>`. New `WORKTREE_CREATED` rows store the `**Worktree path**:` project-relative; `info` resolves it against the project root. Legacy absolute rows remain accepted. The user `cd`s here to inspect a paused Bolt.
- **`branch_name`** — git branch name on which the worktree sits at `bolt-<slug>`, parsed from `**Branch name**:`. Quoted from audit for source-of-truth consistency.
- **`audit_timestamp`** — ISO 8601 timestamp of the matching `WORKTREE_CREATED` block. Useful for the orchestrator to reason about freshness; not currently surfaced in the AUQ prompt.
- **`merge_held`** — boolean reflecting the `Merge-Held` field in the per-Bolt forked state at `<path>/aidlc-docs/aidlc-state.md` (`true` only if the file exists AND the field reads `true`; absence resolves to `false`). The orchestrator reads this on resume to decide whether dispatching `aidlc-bolt complete --merge --slug <slug>` is safe. The held state is set by `aidlc-bolt hold-merge --slug <slug>` before a multi-failure halt-and-ask sequence opens and cleared by `aidlc-bolt release-merge --slug <slug>` once all sibling AUQs resolve.
## Most-recent semantics
`info` returns the **most-recent** `WORKTREE_CREATED` for the slug — meaning the latest by audit-log position (end-to-start walk via `findLatestEvent`). When a slug has been created → discarded → re-created within the same workflow, the second create's path is what `info` returns. This matches the user's mental model: "the live worktree for slug X."
The retry-then-fail scenario (code-gen fails, user picks Retry, code-gen fails again) does not create a new `WORKTREE_CREATED` — Retry re-runs the existing worktree per the SKILL.md per-Bolt loop. So `info`'s output is stable across retry attempts. Pinned by `tests/worktree/t11-halt-and-ask-retry-correlation.sh`.
## Worktree metadata repository provenance
New `.aidlc/worktree-meta.json` files store `gitCommonDirHash`, a 64-character
SHA-256 hex digest of the canonical Git common-directory path. The raw machine
path is not persisted. Merge validation hashes the selected checkout and
worktree common directories and compares the digests.
Migration remains compatible with older metadata carrying plaintext
`gitCommonDir`: readers hash that stored value before comparison. New metadata
must not write both fields.
## Stderr error messages
Three stable messages the orchestrator can route on (though it rarely needs to — exit code is sufficient):
```
error: no WORKTREE_CREATED audit entry for slug <slug> (audit log absent)
error: no WORKTREE_CREATED audit entry for slug <slug>
error: malformed WORKTREE_CREATED block at <timestamp> (missing Worktree path or Branch name field)
```
The third (malformed-block) case is the audit-of-intent reconciliation surface: doctor handles flagging and remediation; `info` just refuses to guess.
## AUQ prompt rendering — long-path fallback
The orchestrator interpolates `path` and `branch_name` into the structured question prompt body, which renders at full terminal width and wraps gracefully (multi-line wrap is supported on macOS Claude Code; verified manually before each release).
If a future surface (Windows PowerShell, mosh, narrow tmux pane) clips long paths in `question`, the documented fallback is to truncate with leading-ellipsis at directory boundaries while preserving the `bolt-<slug>` tail:
```
.../project/.aidlc/worktrees/bolt-onboarding-wizard
```
This fallback is **not currently implemented** — current shipping behaviour assumes graceful wrap. If a regression surfaces, add a `--max-path-display <chars>` flag to `info` and have the orchestrator truncate per the rule above.
## Related files
- Implementation: `.aidlc/tools/aidlc-worktree.ts` (`handleInfo` handler)
- Test: `tests/unit/t72-worktree-info.sh`
- Caller: `.aidlc/skills/aidlc/SKILL.md` (per-Bolt-loop halt-and-ask flow)
- Audit emitter that produces the `WORKTREE_CREATED` entries `info` reads: `aidlc-worktree.ts` `handleCreate` (also in this file at `~line 154`)
- Audit-format spec: `.aidlc/knowledge/aidlc-shared/audit-format.md` `WORKTREE_CREATED` row