# System Architecture Document

## AOLL — Always-On Lead Library
### Platform Architecture & Infrastructure Design

| Field | Value |
|-------|-------|
| **Document Version** | 1.0 |
| **Status** | Draft — For Engineering & Infrastructure Review |
| **Author** | Solutions Architecture |
| **Last Updated** | July 27, 2026 |
| **Companion Documents** | PRD-AOLL.md, TECHNICAL_SPECIFICATION-AOLL.md |
| **Classification** | Internal — Engineering |

---

## Table of Contents

1. [Architecture Overview](#1-architecture-overview)
2. [Architecture Principles](#2-architecture-principles)
3. [High-Level System Context](#3-high-level-system-context)
4. [Logical Architecture](#4-logical-architecture)
5. [Data Architecture](#5-data-architecture)
6. [Application Architecture](#6-application-architecture)
7. [Integration Architecture](#7-integration-architecture)
8. [AI Architecture](#8-ai-architecture)
9. [Security Architecture](#9-security-architecture)
10. [Deployment Architecture](#10-deployment-architecture)
11. [Scalability & Performance](#11-scalability--performance)
12. [Disaster Recovery & Business Continuity](#12-disaster-recovery--business-continuity)
13. [Architecture Decision Records (ADRs)](#13-architecture-decision-records-adrs)
14. [Infeasible Components & Architectural Workarounds](#14-infeasible-components--architectural-workarounds)
15. [Assumptions](#15-assumptions)
16. [Future State Architecture (v2+)](#16-future-state-architecture-v2)

---

## 1. Architecture Overview

AOLL is a **B2B sales intelligence SaaS platform** architected as a **modular system** comprising:

1. A **web application** (existing CodeIgniter 4 monolith) serving UI and authenticated APIs
2. A **product data store** (PostgreSQL) holding canonical company, contact, and signal entities
3. A **search plane** (OpenSearch) optimized for faceted prospecting queries
4. An **async processing plane** (Redis + workers) for alerts, workflows, exports, and indexing
5. A **data pipeline** (batch ETL) for ingestion, entity resolution, and enrichment
6. **External integrations** (Stripe, CRMs, LLM provider, data vendors)

The architecture prioritizes **time-to-market** over premature distribution. Phase 2 introduces the search and data layers while preserving the existing web codebase. Services are extracted only at clear scalability or team boundaries.

### 1.1 Current vs. Target Architecture

| Dimension | Current (CHGREL-1255) | Target (Phase 5) |
|-----------|----------------------|------------------|
| Application | Monolithic CodeIgniter | Modular monolith + search + workers |
| Product data | Mock JSON | PostgreSQL + OpenSearch |
| Auth | MySQL sessions | MySQL/Redis sessions |
| Billing | None | Stripe |
| CRM | UI only | OAuth + one-way sync |
| AI | None | LLM service with RAG |
| Data pipeline | None | Batch ETL + entity resolution |
| API | Mock reads / 501 writes | Full REST `/api/v1/` |

---

## 2. Architecture Principles

| # | Principle | Implication |
|---|-----------|-------------|
| P1 | **Data is the product** | Invest disproportionately in pipeline quality, entity resolution, and search — not just UI |
| P2 | **Separate read and write paths** | Search queries never hit PostgreSQL full-table scans; OpenSearch serves reads |
| P3 | **Stable entity identity** | Every company/contact has an immutable UUID from resolution; all joins use `entity_id` |
| P4 | **Async by default for side effects** | Exports, alerts, workflow runs, emails, and index updates are queued |
| P5 | **Fail closed on entitlements** | Missing or expired subscription → deny metered actions, not silent allow |
| P6 | **Honest data labeling** | Estimated media spend and proxy intent are architecturally tagged at source (`media_spend_source`, `intent.source`) |
| P7 | **Evolutionary architecture** | Extract services when metrics justify — not before |
| P8 | **Vendor abstraction** | Data vendors and LLM providers swappable via adapter interfaces |

---

## 3. High-Level System Context

### 3.1 System Context Diagram

```mermaid
flowchart TB
    subgraph users["Users"]
        sales_user["Sales User"]
        admin_user["Workspace Admin"]
        data_ops["Data Operations"]
    end

    aoll["AOLL Platform"]

    subgraph external["External Systems"]
        stripe["Stripe Billing"]
        hubspot["HubSpot CRM"]
        salesforce["Salesforce CRM"]
        llm["LLM Provider"]
        data_vendor["Data Vendors optional"]
        email["Email Provider"]
        google["Google Sign-In"]
    end

    sales_user -->|"Search and export"| aoll
    admin_user -->|"Manage team and billing"| aoll
    data_ops -->|"Upload CSV Excel monitor scraper"| aoll

    aoll -->|"Subscriptions"| stripe
    aoll -->|"Push records"| hubspot
    aoll -->|"Push records"| salesforce
    aoll -->|"AI generation"| llm
    aoll -->|"Optional vendor data"| data_vendor
    aoll -->|"Send email"| email
    aoll -->|"Sign in"| google
```

> **Auth scope (v1):** Users sign in with **email + password** or **Google OAuth** only. Microsoft/Office 365 SSO is **out of scope**.

> **Data ingestion (internal to AOLL):** An existing **web scraper** runs on a schedule and dumps records into staging/DB. **Admin CSV/Excel upload** accepts bulk files through an internal admin UI. Both feed the same normalize → entity resolution → search index pipeline. See §5.2 and M17 in the PRD.

### 3.2 Primary User Journeys (Architecturally Significant)

| Journey | Systems Touched | Latency Sensitivity |
|---------|-----------------|---------------------|
| Advanced Search → Company Profile | Web → API → OpenSearch → PostgreSQL | High (< 2s) |
| Save Search → Alert Email | Web → PostgreSQL → Worker → OpenSearch diff → Email | Low (hours) |
| Workflow → CRM Export | Worker → PostgreSQL → HubSpot/SF API | Medium (minutes) |
| AI Outreach Generation | Web → PostgreSQL → LLM API | Medium (< 10s) |
| Signup → Stripe Checkout | Web → Stripe → Webhook → PostgreSQL | Medium |

---

## 4. Logical Architecture

### 4.1 Container Diagram

```mermaid
flowchart TB
    subgraph clientLayer["Client Layer"]
        Browser["Web Browser"]
    end

    subgraph edgeLayer["Edge Layer"]
        LB["Load Balancer and CDN"]
        WAF["WAF"]
    end

    subgraph appLayer["Application Layer"]
        CI["CodeIgniter 4 Web App"]
        API["REST API v1"]
    end

    subgraph serviceLayer["Service Layer"]
        SearchSvc["Search Query Service"]
        EntitleSvc["Entitlement Service"]
        AISvc["AI Orchestration Service"]
        CRMsvc["CRM Connector Service"]
    end

    subgraph asyncLayer["Async Processing Layer"]
        Redis["Redis Queue and Cache"]
        Workers["Background Workers"]
    end

    subgraph dataLayer["Data Layer"]
        PG["PostgreSQL Product DB"]
        MySQL["MySQL CMS Users"]
        OS["OpenSearch Cluster"]
        S3["Object Storage"]
    end

    subgraph pipelineLayer["Data Pipeline Layer"]
        Scraper["Web Scraper"]
        Upload["CSV Excel Import"]
        Airflow["Scheduler"]
        ETL["ETL Workers"]
        Resolver["Entity Resolution Engine"]
    end

    subgraph externalLayer["External Systems"]
        Stripe["Stripe"]
        LLM["LLM API"]
        Vendors["Data Vendors"]
        HS["HubSpot"]
        SF["Salesforce"]
    end

    Browser --> LB --> WAF --> CI
    CI --> API
    API --> SearchSvc --> OS
    API --> EntitleSvc --> PG
    API --> AISvc --> LLM
    API --> CRMsvc --> HS
    API --> CRMsvc --> SF
    API --> Redis
    Workers --> Redis
    Workers --> PG
    Workers --> OS
    Workers --> S3
    Workers --> HS
    Workers --> SF
    CI --> MySQL
    API --> PG
    Scraper --> ETL
    Upload --> ETL
    Airflow --> ETL --> S3
    ETL --> Resolver --> PG
    ETL --> OS
    Vendors --> ETL
    Stripe --> CI
    PG -.->|batch index| OS
```

### 4.2 Layer Responsibilities

| Layer | Responsibility | Technology |
|-------|----------------|------------|
| **Client** | Dashboard UI, marketing site | HTML/CSS/JS (server-rendered CI4 views) |
| **Edge** | TLS termination, static assets, DDoS protection | CloudFront / nginx |
| **Application** | HTTP routing, session auth, API gateway, view rendering | CodeIgniter 4 |
| **Services** | Domain logic: search DSL, entitlements, AI prompts, CRM adapters | PHP service classes |
| **Async** | Job queue, scheduled tasks, retries | Redis + workers |
| **Data** | Persistent storage, search index, file storage | PostgreSQL, OpenSearch, S3 |
| **Pipeline** | Ingestion, resolution, enrichment, index rebuild | Python ETL |

---

## 5. Data Architecture

### 5.1 Data Domain Model

```mermaid
erDiagram
    WORKSPACE ||--o{ USER : has
    WORKSPACE ||--o| SUBSCRIPTION : has
    WORKSPACE ||--o{ SAVED_SEARCH : owns
    WORKSPACE ||--o{ WORKFLOW : owns
    WORKSPACE ||--o{ CONNECTOR : owns

    COMPANY ||--o{ CONTACT : employs
    COMPANY ||--o{ SCOOP : generates
    COMPANY ||--o{ INTENT_SIGNAL : shows
    COMPANY ||--o{ COMPANY_TECHNOLOGY : uses

    BRAND ||--o{ BRAND_AGENCY_REL : works_with
    AGENCY ||--o{ BRAND_AGENCY_REL : represents

    TOPIC ||--o{ INTENT_SIGNAL : categorizes
    USER ||--o{ USER_TOPIC : tracks
    TOPIC ||--o{ USER_TOPIC : assigned

    SAVED_SEARCH ||--o{ ALERT : produces
    WORKFLOW ||--o{ WORKFLOW_RUN : executes
```

### 5.2 Data Flow — Ingestion to Serving

```mermaid
flowchart LR
    subgraph sources["Sources"]
        V0["Web Scraper"]
        V1["CSV Excel Upload"]
        V2["Data Vendor API"]
        V3["News and RSS"]
        V4["Ad Libraries"]
        V5["Intent Feed"]
    end

    subgraph landing["Landing"]
        S3Raw["S3 Raw Zone"]
    end

    subgraph processing["Processing"]
        Norm["Normalizer"]
        NER["Entity Extractor"]
        ER["Entity Resolver"]
        Enrich["Enrichment Engine"]
        Score["Signal Scorer"]
    end

    subgraph storage["Storage"]
        PG["PostgreSQL"]
        OS["OpenSearch"]
    end

    subgraph serving["Serving"]
        API["AOLL API"]
        UI["Dashboard"]
    end

    V0 --> S3Raw
    V1 --> S3Raw
    V2 --> S3Raw
    V3 --> S3Raw
    V4 --> S3Raw
    V5 --> S3Raw
    S3Raw --> Norm --> NER --> ER --> Enrich --> Score --> PG
    PG -->|index sync| OS
    OS --> API --> UI
    PG --> API
```

### 5.3 Entity Resolution Architecture

Entity resolution is the **most critical data subsystem**. Without it, search results contain duplicates, CRM exports create conflicting records, and signals attach to wrong companies.

```mermaid
flowchart TD
    Input["Raw Record"]
    Normalize["Normalize fields"]
    T1{"Tier 1: exact domain match?"}
    T2{"Tier 2: fuzzy name match gte 0.92?"}
    T3["Tier 3: LLM adjudication"]
    Existing["Return existing entity_id"]
    New["Create new entity_id"]
    Output["Resolved Entity"]

    Input --> Normalize --> T1
    T1 -->|Yes| Existing
    T1 -->|No| T2
    T2 -->|Yes| Existing
    T2 -->|Ambiguous| T3
    T2 -->|No match| New
    T3 -->|Merge| Existing
    T3 -->|Distinct| New
    Existing --> Output
    New --> Output
```

**Architectural constraints:**
- Resolution runs in **pipeline (batch)**, not synchronously on user search (except domain lookup cache)
- All downstream systems **must** reference `entity_id`, never raw company name alone
- Merge operations are **audited** in `entity_merge_log` table

### 5.4 Data Classification

| Class | Examples | Handling |
|-------|----------|----------|
| **Public firmographics** | Company name, industry, website | Indexable; no access gate beyond plan |
| **Licensed PII** | Contact emails, phone numbers | Metered; usage tracked; export logged |
| **Estimated/proxy data** | Media spend estimate, proxy intent | Must carry `source` metadata; UI disclaimer |
| **User-generated** | Saved searches, feed preferences | Workspace-scoped; deleted with account |
| **Credentials** | CRM OAuth tokens | Encrypted at rest; never logged |

### 5.5 Data Freshness SLAs

| Data Type | Target Freshness | Mechanism |
|-----------|------------------|-----------|
| Firmographics (vendor) | Weekly refresh | Vendor sync job |
| Scoops | ≤ 48 hours from publication | Daily news pipeline |
| Intent (licensed) | Daily | Bombora batch import |
| Intent (proxy) | Daily | Composite recalculation job |
| Media spend estimates | Weekly | Ad library scan job |
| Search index | ≤ 15 min (target) / ≤ 24h (v1 acceptable) | Index sync job |
| CRM export | Real-time on user action | Async worker |

---

## 6. Application Architecture

### 6.1 CodeIgniter Module Structure (Target)

```
aoll.com/
├── app/
│   ├── Controllers/
│   │   ├── Dashboard/          # UI controllers (existing)
│   │   └── Api/V1/             # REST API controllers (new)
│   │       ├── SearchController.php
│   │       ├── CompanyController.php
│   │       ├── IntentController.php
│   │       ├── FeedController.php
│   │       ├── SavedSearchController.php
│   │       ├── AlertController.php
│   │       ├── WorkflowController.php
│   │       ├── ConnectorController.php
│   │       ├── BillingController.php
│   │       ├── AiController.php
│   │       └── AdminController.php
│   ├── Services/
│   │   ├── Search/             # OpenSearch query builder
│   │   ├── Entitlement/        # Plan limits, usage metering
│   │   ├── Feed/               # Feed ranking
│   │   ├── Alert/              # Alert diff engine
│   │   ├── Workflow/           # Workflow executor
│   │   ├── Crm/                # HubSpot, Salesforce adapters
│   │   ├── Ai/                 # LLM orchestration
│   │   ├── Export/             # CSV generation
│   │   └── Billing/            # Stripe integration
│   ├── Models/                 # PostgreSQL Eloquent/CI models
│   ├── Filters/
│   │   ├── AuthFilter.php
│   │   ├── EntitlementFilter.php
│   │   └── AdminFilter.php
│   └── Jobs/                   # Job definitions (queued to Redis)
├── workers/
│   ├── alert_worker.php
│   ├── workflow_worker.php
│   ├── export_worker.php
│   └── index_sync_worker.php
└── pipeline/                   # Python ETL (separate repo or subdir)
    ├── ingest/
    ├── resolve/
    ├── enrich/
    └── index/
```

### 6.2 Request Flow — Advanced Search

```mermaid
sequenceDiagram
    participant U as User Browser
    participant CI as CodeIgniter API
    participant E as Entitlement Service
    participant S as Search Service
    participant OS as OpenSearch
    participant PG as PostgreSQL

    U->>CI: GET search companies
    CI->>CI: Validate session
    CI->>E: check search quota
    E->>PG: get usage count and plan tier
    alt Limit exceeded
        E-->>CI: 402 Payment Required
        CI-->>U: Upgrade prompt
    else Allowed
        E->>PG: increment usage
        CI->>S: buildQuery filters
        S->>OS: Execute DSL query
        OS-->>S: Hits and aggregations
        S-->>CI: Normalized results
        CI-->>U: JSON response
    end
```

### 6.3 Request Flow — Workflow Execution

```mermaid
sequenceDiagram
    participant Cron as Scheduler
    participant W as Workflow Worker
    participant PG as PostgreSQL
    participant OS as OpenSearch
    participant CRM as HubSpot or SF
    participant Email as Email Service

    Cron->>W: Trigger due workflows
    W->>PG: Load workflow config
    W->>OS: Execute trigger query
    OS-->>W: Company entity IDs
    
    loop Each action
        alt Export to CRM
            W->>PG: Load contacts for entities
            W->>CRM: Push records batched
            alt CRM failure
                CRM-->>W: Error
                W->>PG: Set connector status error
                W->>Email: Queue admin notification
            end
        else Discover Contacts
            W->>PG: Query contacts by role
        else Webhook
            W->>W: POST payload to URL
        end
    end
    
    W->>PG: Update last_run_at, next_run_at
```

### 6.4 Feed Personalization Architecture

The home feed is a **ranked composite** of scoops and intent signals, not a separate data store.

```
feed_items = UNION(
    SELECT scoop.*, 'scoop' AS item_type FROM scoops WHERE company matches ICP,
    SELECT intent.*, 'intent' AS item_type FROM intent_signals WHERE topic IN user_topics
)
ORDER BY ranking_score DESC
LIMIT 50
```

**ICP match** computed from `feed_preferences` table (industries, locations, sizes, media spend tiers).

Cached in Redis per user (`feed:{user_id}`, TTL 15 min) invalidated on preference change.

---

## 7. Integration Architecture

### 7.1 Integration Map

```mermaid
flowchart LR
    subgraph aollSystem["AOLL"]
        App["Web App"]
        WH["Webhook Handler"]
    end

    subgraph billingZone["Billing"]
        Stripe["Stripe"]
    end

    subgraph crmZone["CRM"]
        HS["HubSpot"]
        SF["Salesforce"]
    end

    subgraph identityZone["Identity"]
        Google["Google Sign-In"]
    end

    subgraph dataZone["Data Providers"]
        Vendor["Data Vendor"]
        Bombora["Bombora optional"]
    end

    subgraph commsZone["Comms"]
        Email["SendGrid"]
    end

    App --> Stripe
    Stripe --> WH
    App --> HS
    App --> SF
    App --> Google
    Vendor --> App
    Bombora --> App
    App --> Email
```

### 7.2 CRM Connector Pattern (Adapter)

All CRM integrations implement a common interface:

```php
interface CrmConnectorInterface {
    public function getAuthorizationUrl(): string;
    public function handleCallback(string $code): ConnectorCredentials;
    public function refreshToken(ConnectorCredentials $creds): ConnectorCredentials;
    public function pushCompanies(array $companies): PushResult;
    public function pushContacts(array $contacts): PushResult;
    public function healthCheck(): HealthStatus;
}
```

| Concern | Pattern |
|---------|---------|
| Token storage | Encrypted in `connectors` table; refresh 24h before expiry |
| Rate limiting | Token bucket per connector; exponential backoff |
| Idempotency | Upsert by email (contact) / domain (company) |
| Failure handling | 3 retries → mark connector error → email admin |
| Field mapping | JSON config per workspace; defaults per provider |

**v1 limitation:** One-way push only. No inbound CRM webhooks.

### 7.3 Stripe Integration Architecture

```mermaid
sequenceDiagram
    participant U as User
    participant App as AOLL
    participant S as Stripe
    participant PG as PostgreSQL

    U->>App: Click Upgrade to Professional
    App->>S: Create Checkout Session
    S-->>App: session_url
    App-->>U: Redirect to Stripe
    U->>S: Complete payment
    S->>App: Webhook payment completed
    App->>PG: Update workspace plan tier
    App->>App: Invalidate entitlement cache
    App-->>U: Email plan upgrade
```

**Webhook security:** Verify `Stripe-Signature` header; store processed event IDs to prevent replay.

---

## 8. AI Architecture

### 8.1 AI System Context

AI capabilities are **downstream of data quality**. The architecture grounds all LLM calls in structured company context retrieved from PostgreSQL — never raw user input alone.

```mermaid
flowchart TB
    Req["API Request"] --> CB["Context Builder"]
    PG2["PostgreSQL"] --> CB
    CB --> Guard["Input Validator"]
    Guard --> Template["Prompt Template"]
    Template --> Model["LLM Model"]
    Model --> Validate["Output Validator"]
    Validate --> Store["Generation Log"]
    Store --> Response["API Response"]
```

### 8.2 AI Feature Matrix

| Feature | v1 Approach | v2 Evolution |
|---------|-------------|--------------|
| Opportunity Score | Rule-based weighted formula | ML classifier on conversion data |
| Opportunity Summary | LLM with RAG context | Fine-tuned model |
| Budget Estimate | Heuristic ranges | Regression model |
| Email Outreach | LLM prompt templates | A/B tested prompts; user feedback loop |
| LinkedIn Messages | LLM prompt templates | Same |
| Call Scripts | LLM prompt templates | Same |

### 8.3 AI Guardrails (Architectural)

| Guardrail | Implementation |
|-----------|----------------|
| No fabricated facts | System prompt + output validation against context fields |
| Cost control | Per-workspace monthly generation quota |
| Audit trail | `ai_generations` table: prompt_hash, model, tokens, user_id, timestamp |
| PII in prompts | Contact email included only when user explicitly requests outreach for that contact |
| Failure mode | LLM timeout/error → graceful 503 with retry message; never return empty silently |

### 8.4 What AI Cannot Reliably Do (Honest Assessment)

| Claim | Feasible? | Architectural Response |
|-------|-----------|------------------------|
| "This company will buy in 30 days" | ❌ | Frame as "opportunity score" with contributing factors |
| Accurate budget without spend data | ❌ | Show ranges with low confidence; label as estimate |
| 100% personalized outreach without hallucination | ⚠️ | RAG grounding + disclaimer: "Review before sending" |
| Replace human research entirely | ❌ | Position as draft accelerator, not autonomous agent |

---

## 9. Security Architecture

### 9.1 Security Zones

```mermaid
flowchart TB
    Marketing["Marketing Site"] --> Dashboard["Dashboard"]
    Login["Login and Signup"] --> Dashboard
    Dashboard --> PG["PostgreSQL"]
    Dashboard --> OS["OpenSearch"]
    Dashboard --> Redis["Redis"]
    Dashboard --> Workers["Workers"]
    ETL["ETL Jobs"] --> PG
    ETL --> S3["S3 Raw Data"]
```

### 9.2 Authentication & Authorization Model

| Layer | Mechanism |
|-------|-----------|
| Authentication | Session cookie (web) / API key (Enterprise) |
| Session store | Redis with TTL |
| Authorization — role | Admin vs Non-Admin (RBAC) |
| Authorization — plan | Entitlement service checks feature flags + usage limits |
| Admin actions | Billing, user management, connector setup — Admin only |
| API key scope | Enterprise workspace; revocable; rate limited |

### 9.3 Threat Model Summary

| Threat | Mitigation |
|--------|------------|
| Credential stuffing | Rate limit login; generic error messages |
| Session hijacking | HTTPS only; HttpOnly Secure cookies |
| Mass contact scraping | Usage limits; rate limiting; export logging |
| CRM token theft | Encryption at rest; no token in logs |
| Stripe webhook forgery | Signature verification |
| LLM prompt injection | Context builder uses DB fields only; user input sanitized |
| SQL injection | Parameterized queries |
| XSS | Output encoding in views |

---

## 10. Deployment Architecture

### 10.1 Production Topology (Target)

```mermaid
flowchart TB
    Users["Users"] --> CF["CloudFront CDN"]
    CF --> ALB["Application Load Balancer"]
    ALB --> NG1["PHP Instance 1"]
    ALB --> NG2["PHP Instance 2"]
    NG1 --> RDS["PostgreSQL Multi-AZ"]
    NG2 --> RDS
    NG1 --> OSCluster["OpenSearch Cluster"]
    NG1 --> RedisCluster["ElastiCache Redis"]
    NG2 --> OSCluster
    Worker1["Worker Instance"] --> RedisCluster
    Worker1 --> RDS
    Worker1 --> OSCluster
    AirflowNode["Airflow Scheduler"] --> ETLNode["ETL Workers"]
    ETLNode --> S3Bucket["S3 Bucket"]
    ETLNode --> RDS
    ETLNode --> OSCluster
```

### 10.2 Environment Strategy

| Environment | Infrastructure | Data Volume |
|-------------|----------------|-------------|
| **Development** | Docker Compose locally | 1K seed companies |
| **Staging** | Scaled-down prod mirror | 10K anonymized records |
| **Production** | Full topology | 100K+ companies (growing) |

### 10.3 CI/CD Pipeline

```mermaid
flowchart LR
    Commit["Git Push"] --> Build["Build and Test"]
    Build --> Staging["Deploy Staging"]
    Staging --> QA["QA Approval"]
    QA --> Prod["Deploy Production"]
    Build --> UnitTests["Unit Tests"]
    Staging --> SmokeTests["Smoke Tests"]
```

### 10.4 Infrastructure Sizing (Initial Production Estimate)

| Component | Initial Size | Notes |
|-----------|--------------|-------|
| PHP-FPM instances | 2 × t3.medium | Auto-scale at CPU > 70% |
| Worker instance | 1 × t3.small | Scale with queue depth |
| PostgreSQL | db.r6g.large, 100GB | Multi-AZ |
| OpenSearch | 3 × r6g.large.search | 100GB each |
| Redis | cache.r6g.large | Single node v1; cluster v2 |
| S3 | Standard tier | Lifecycle policy on raw data (90d) |

*Sizing to be validated by load testing before production launch.*

---

## 11. Scalability & Performance

### 11.1 Scaling Strategy by Component

| Component | Scale Trigger | Strategy |
|-----------|---------------|----------|
| Web app (PHP) | CPU > 70% or latency P95 > 1s | Horizontal: add instances behind ALB |
| OpenSearch | Query latency P95 > 2s or CPU > 80% | Add data nodes; optimize queries |
| PostgreSQL | Connection pool exhaustion | Read replicas for analytics; PgBouncer |
| Redis queue | Queue depth > 1000 for > 5 min | Add worker instances |
| ETL pipeline | Job duration > SLA | Parallelize by source; add workers |
| LLM API | Rate limit errors | Queue AI requests; backoff |

### 11.2 Caching Strategy

| Cache | Key | TTL | Invalidation |
|-------|-----|-----|--------------|
| Feed | `feed:{user_id}` | 15 min | Preference change |
| Company profile | `company:{entity_id}` | 1 hour | Entity update webhook |
| Plan entitlements | `entitlements:{workspace_id}` | 5 min | Stripe webhook |
| Search facets | `facets:{entity_type}` | 1 hour | Index rebuild |
| OAuth tokens | — | Never cached in Redis | DB only, encrypted |

### 11.3 Performance Budget

| User Action | P50 | P95 | P99 |
|-------------|-----|-----|-----|
| Search (filtered) | 400ms | 2000ms | 4000ms |
| Company profile load | 200ms | 800ms | 1500ms |
| Feed load | 300ms | 1500ms | 3000ms |
| AI outreach generation | 3s | 8s | 15s |
| CSV export (1K rows) | — | 30s (async) | 60s |
| CRM push (100 contacts) | — | 60s (async) | 120s |

---

## 12. Disaster Recovery & Business Continuity

| Scenario | RTO | RPO | Strategy |
|----------|-----|-----|----------|
| AZ failure | 1 hour | 5 min | Multi-AZ RDS; OpenSearch zone awareness |
| Database corruption | 4 hours | 24 hours | Daily snapshots; point-in-time recovery |
| OpenSearch cluster failure | 2 hours | 24 hours | Rebuild index from PostgreSQL |
| Stripe outage | N/A | N/A | Graceful degradation; existing subscriptions remain active |
| Data vendor outage | N/A | 7 days | Serve stale data with `last_verified_at` displayed |
| LLM provider outage | N/A | N/A | AI features return 503; core search unaffected |

**Backup schedule:**
- PostgreSQL: automated daily snapshots, 35-day retention
- OpenSearch: index rebuild from PostgreSQL (preferred over snapshot restore)
- S3 raw data: versioning enabled; 90-day lifecycle

---

## 13. Architecture Decision Records (ADRs)

### ADR-001: Retain CodeIgniter 4 as Web Framework

| | |
|---|---|
| **Status** | Accepted |
| **Context** | CHGREL-1255 delivered working auth, UI, and routing in CI4 |
| **Decision** | Continue with CodeIgniter 4 for web + API; do not rewrite in Node/React SSR for v1 |
| **Consequences** | Faster delivery; PHP talent required; SPA migration deferred |
| **Alternatives rejected** | Full Next.js rewrite (6+ month delay); Laravel migration (no benefit) |

### ADR-002: Introduce OpenSearch as Separate Search Plane

| | |
|---|---|
| **Status** | Accepted |
| **Context** | Advanced search requires 300+ filter combinations at sub-2s latency |
| **Decision** | OpenSearch cluster; PostgreSQL remains system of record |
| **Consequences** | Index sync complexity; operational overhead; proven at scale |
| **Alternatives rejected** | PostgreSQL full-text only (insufficient for faceted search); Algolia (cost at scale) |

### ADR-003: Batch Data Pipeline (Not Streaming) for v1

| | |
|---|---|
| **Status** | Accepted |
| **Context** | Real-time streaming requires Kafka + Flink; team size and timeline constrained |
| **Decision** | Daily batch ETL with Airflow/cron; daily intent/scoop refresh |
| **Consequences** | Data freshness ≤ 24–48h; acceptable per PRD assumptions |
| **Alternatives rejected** | Kafka streaming (over-engineered for v1) |

### ADR-004: Proxy Intent as v1 Workaround

| | |
|---|---|
| **Status** | Accepted (pending Bombora budget approval) |
| **Context** | True content-consumption intent requires Bombora license ($10K–20K/yr) |
| **Decision** | Ship proxy composite intent in Phase 3; label clearly in UI |
| **Consequences** | Weaker intent story vs ZoomInfo until licensed; honest UX required |
| **Alternatives rejected** | Fake intent scores (trust-destroying); delay launch until Bombora (timeline risk) |

### ADR-005: One-Way CRM Export Only

| | |
|---|---|
| **Status** | Accepted |
| **Context** | Bi-directional sync requires conflict resolution, duplicate handling, field mapping UI |
| **Decision** | Push-only export to HubSpot/Salesforce in v1 |
| **Consequences** | Users manage CRM duplicates manually; 3–6 months saved |
| **Alternatives rejected** | Bi-directional v1 (scope explosion) |

### ADR-006: License Base Data, Enrich Proprietary

| | |
|---|---|
| **Status** | Accepted |
| **Context** | Building 100M+ contact database from scratch is not feasible |
| **Decision** | License company/contact base from Apollo or Clearbit; build proprietary ad-spend and brand-agency enrichment |
| **Consequences** | Ongoing vendor cost ($50K–200K/yr); faster time to credible data |
| **Alternatives rejected** | Full proprietary crawl (12–18 months, legal risk) |

---

## 14. Infeasible Components & Architectural Workarounds

This section documents capabilities that **cannot be built as specified** within v1 constraints, with approved architectural workarounds.

| # | Requested Capability | Why Infeasible | Architectural Workaround | User-Facing Label |
|---|---------------------|----------------|--------------------------|-------------------|
| 1 | ZoomInfo-scale database (100M+ companies) | Requires $100M+ and years of data ops | License vendor data; vertical focus (100K–500K) | Coverage disclaimer on search |
| 2 | Winmo-verified ad spend ($100B+ tracked) | Proprietary aggregation over decades | Meta/Google Ad Library → estimation model | **"Estimated media spend"** |
| 3 | Bombora content-consumption intent | Requires publisher network partnership | Proxy composite from scoops + hiring + funding + ad signals | **"Activity-based intent score"** |
| 4 | Real-time intent alerts | Requires streaming pipeline | Daily/weekly batch diff on saved searches | "Updated daily" badge |
| 5 | Real-time workflow event triggers | Requires event bus + stream processing | Scheduled daily/weekly workflow runs | Schedule picker in UI |
| 6 | Website visitor ID (WebSights) | Requires pixel + identity graph at scale | **Not built** — out of scope v1 | N/A |
| 7 | Org charts / reporting lines | LinkedIn ToS; manual research at scale | Flat contact list by seniority | "Decision makers" section |
| 8 | Bi-directional CRM sync | Conflict resolution complexity | One-way push with upsert keys | "Export to CRM" (not "Sync") |
| 9 | 95% contact email accuracy | Requires human verification team | Verification status enum; periodic SMTP check sample | Verified / Likely / Unverified badges |
| 10 | LinkedIn data scraping | Legal and ToS violation | Licensed contact data from vendor | N/A |
| 11 | MCP / headless AI agent API | ZoomInfo 2026 feature; premature for AOLL | Enterprise REST API in Phase 5 | API docs for Enterprise |
| 12 | Brand–agency mapping at Winmo completeness | Relationship data is manually curated | NLP extraction from press + curated seed dataset (5K brands) | "Known relationships" with source |
| 13 | Microsoft SSO (Day 1) | Incremental OAuth scope | Google OAuth first (done); Microsoft Phase 3 | N/A |
| 14 | Pipedrive integration | Resource constraint vs HubSpot/SF | Defer to post-Phase 4 | Roadmap item |

---

## 15. Assumptions

### 15.1 Business Assumptions

| ID | Assumption |
|----|------------|
| BA1 | AOLL launches US-market first; GDPR/CCPA compliance built but EU expansion deferred |
| BA2 | Primary ICP is media/advertising vertical; general B2B is secondary |
| BA3 | Self-serve pricing is a competitive advantage worth maintaining |
| BA4 | Users accept daily data refresh (not real-time) for v1 |
| BA5 | Data licensing budget of $50K–$200K/year is approved for Phase 2 |

### 15.2 Technical Assumptions

| ID | Assumption |
|----|------------|
| TA1 | Existing CodeIgniter 4 codebase (`aoll.com`) is the deployment unit for web + API |
| TA2 | AWS (or equivalent cloud) is available for managed PostgreSQL, OpenSearch, Redis, S3 |
| TA3 | Engineering team can operate 2 additional infrastructure components (OpenSearch, PostgreSQL) |
| TA4 | Python is acceptable for ETL pipeline (may be separate repo) |
| TA5 | Single-region (us-east-1) deployment sufficient for v1 |
| TA6 | Horizontal scaling of PHP-FPM behind ALB handles 10K MAU |
| TA7 | Async workers process 100K alert diffs/night within 4-hour window |

### 15.3 Data Assumptions

| ID | Assumption |
|----|------------|
| DA1 | Licensed vendor data provides ≥80% coverage for target vertical companies |
| DA2 | Public ad library APIs remain accessible for media spend estimation |
| DA3 | News/RSS sources provide sufficient scoop volume (≥100/day relevant) |
| DA4 | Entity resolution achieves ≤5% duplicate rate at 0.92 fuzzy threshold |
| DA5 | Proxy intent correlates sufficiently with sales outcomes to retain users until Bombora |

---

## 16. Future State Architecture (v2+)

Components intentionally deferred beyond Phase 5:

```mermaid
flowchart TB
    Core["Core Platform Phase 5"]
    Chrome["Chrome Extension"]
    MCP["MCP Agent API"]
    Stream["Kafka Event Bus"]
    RTIntent["Real-time Intent"]
    WebSights["Visitor ID Pixel"]
    OrgChart["Org Chart Service"]
    BiCRM["Bi-directional CRM"]
    MLScore["ML Scoring Platform"]
    MultiRegion["Multi-region Deploy"]

    Core --> Chrome
    Core --> MCP
    Core --> Stream
    Stream --> RTIntent
    Core --> WebSights
    Core --> OrgChart
    Core --> BiCRM
    Core --> MLScore
    Core --> MultiRegion
```

| Component | Trigger to Build | Estimated Effort |
|-----------|------------------|------------------|
| Chrome Extension | Core data credible + 1K paid users | 3 months |
| Real-time intent (Kafka) | Bombora license + 10K MAU | 4 months |
| WebSights pixel | Identity graph maturity | 6+ months |
| Org charts | LinkedIn partnership or manual team | 6+ months |
| MCP/Agent API | Enterprise customer demand | 2 months |
| ML opportunity scoring | ≥10K labeled conversion events | 3 months |
| Multi-region | EU customer demand | 4 months |

---

## Document History

| Version | Date | Author | Changes |
|---------|------|--------|---------|
| 1.0 | 2026-07-27 | Solutions Architecture | Initial architecture document |

---

## Related Documents

| Document | Path | Purpose |
|----------|------|---------|
| Product Requirements Document | `docs/PRD-AOLL.md` | What to build |
| Technical Specification | `docs/TECHNICAL_SPECIFICATION-AOLL.md` | How to build (detailed) |
| Jira Epic | AA-25 | Ticket-level requirements |
| Initial Release Notes | CHGREL-1255 | Current baseline |
| Design System | design.admedia.com/aoll/ | UI reference |

---

*This architecture document is the authoritative reference for AOLL system design. Changes require architecture review and ADR updates.*
