# WebCrawlers SEO Automation Platform

An AI-powered SEO automation platform built on CodeIgniter 4. It gives logged-in users a full dashboard to manage website projects, run technical audits, track rankings, generate AI content, and measure performance — all driven by a versioned REST API that the dashboard UI, and eventually mobile apps and public integrations, consume.

---

## Table of Contents

1. [How This Repo Works](#how-this-repo-works)
2. [Developer Setup](#developer-setup)
3. [URL Structure](#url-structure)
4. [Environment Detection](#environment-detection)
5. [Configuration & Secrets](#configuration--secrets)
6. [Database](#database)
7. [Background Jobs (Queue)](#background-jobs-queue)
8. [Authentication](#authentication)
9. [Project Architecture](#project-architecture)
10. [Folder Structure](#folder-structure)
11. [Adding a New Dashboard Page](#adding-a-new-dashboard-page)
12. [Adding a New API Endpoint](#adding-a-new-api-endpoint)
13. [Design Assets](#design-assets)
14. [External Integrations](#external-integrations)
15. [Development Phases](#development-phases)

---

## How This Repo Works

The project runs on a **shared dev server** at `/var/www/php82/<username>/webcrawlers_latest.com/`. Multiple developers work on the same server; each developer clones into their own home directory under `/var/www/php82/<username>/`.

The server is already configured with PHP 8.2 and MySQL. There is **no Docker, no local install needed**. You edit files directly on the server (via SSH, SFTP, or an IDE with remote support) and hit the URL in your browser to test.

There is **no `.env` file** — all config lives in tracked PHP files so every developer shares the same settings automatically after a `git pull`.

---

## Developer Setup

### 1. Clone the repo

```bash
cd /var/www/php82/<your-username>/
git clone <repo-url> webcrawlers_latest.com
cd webcrawlers_latest.com
```

### 2. Set folder permissions

```bash
chmod -R 775 writable/
```

### 4. Access the site

| Path | URL |
|---|---|
| Public website | `http://192.168.30.106/php82/<your-username>/webcrawlers_latest.com/public/` |
| Dashboard | `http://192.168.30.106/php82/<your-username>/webcrawlers_latest.com/public/?/dashboard` |
| API | `http://192.168.30.106/php82/<your-username>/webcrawlers_latest.com/public/?/api/v1/dashboard` |

> The `?/` prefix is the query-string router used on the shared dev server (no rewrite rules). Routes work normally; the `indexPage = '?'` in `App.php` handles it automatically.

### 5. Fill in API keys

Open `app/Config/Constants.php` and set the `WC_*` constants — see [Configuration & Secrets](#configuration--secrets).

---

## URL Structure

```
webcrawlers.com/              ← public marketing site  (Pages controller)
webcrawlers.com/dashboard     ← logged-in app UI       (Dashboard controller)
webcrawlers.com/api/v1/       ← REST API               (Api\V1\* controllers)
```

All `/dashboard/*` routes are protected by `AuthFilter` (session check).  
All `/api/v1/*` routes are protected by `ApiAuthFilter` (Bearer token or session).

The full route list is in `app/Config/Routes.php`.

---

## Environment Detection

Environment is detected **automatically** in `public/index.php` before CI4 boots, and sets `$_SERVER['CI_ENVIRONMENT']` accordingly:

| Runtime context | `CI_ENVIRONMENT` | Behaviour |
|---|---|---|
| `/var/www/php82/…` (any dev sandbox) | `development` | Full errors, debug toolbar |
| `staging.webcrawlers.com` | `testing` | Prod-like, toolbar available |
| `webcrawlers.com` / anything else | `production` | Errors suppressed, toolbar off |

`App::__construct()` uses the same logic to set the correct `$baseURL` and `$indexPage` for each context — no manual config needed when switching between environments.

---

## Configuration & Secrets

**No `.env` file is used.** All secrets live in `app/Config/Constants.php` so they are shared across the team via Git.

> Before production release, move sensitive values to the real server environment or a `.env` file that is `.gitignore`d, and restore empty defaults here.

| Constant | Purpose |
|---|---|
| `WC_ENCRYPTION_KEY` | Signs API tokens (HMAC-SHA256). Generate with `php -r "echo bin2hex(random_bytes(32));"` |
| `WC_OPENAI_API_KEY` | OpenAI — content generation, AI assistant |
| `WC_OPENAI_MODEL` | OpenAI model name (default `gpt-4o`) |
| `WC_DATAFORSEO_LOGIN` | DataForSEO account login |
| `WC_DATAFORSEO_PASSWORD` | DataForSEO account password |
| `WC_GOOGLE_CLIENT_ID` | Google OAuth2 client ID |
| `WC_GOOGLE_CLIENT_SECRET` | Google OAuth2 client secret |
| `WC_GOOGLE_REDIRECT_URI` | Google OAuth2 redirect URI |


`Encryption.php` references `WC_ENCRYPTION_KEY` directly so the CI4 encryption service and HMAC token signing always use the same key.

**Database credentials** come from `MSACOMMON_PATH . 'dbConstants.php'` — a shared library on the dev server. `Database.php` reads `MSACOMMON_DB_HOST_CMS`, `MSACOMMON_DB_USER_CMS`, and `MSACOMMON_DB_PASS_CMS` from it automatically.

---

## Database

- **Driver:** MySQL via `MySQLi`
- **Default database:** `cms`
- **Credentials:** pulled from `msacommon` shared library (no manual config needed on the dev server)
- **Migrations:** `app/Database/Migrations/`
- **Seeds:** `app/Database/Seeds/`

### Core tables to be created via migrations

```
users                   projects               project_settings
project_integrations    crawl_runs             crawl_pages
crawl_errors            crawl_changes          issues
audit_reports           keywords               keyword_history
keyword_rankings        tracked_keywords       ranking_competitors
competitor_rankings     ranking_alerts         keyword_searches
saved_keywords          keyword_gap_runs       content_analyses
content_briefs          content_drafts         content_change_jobs
content_change_audit    leads                  queue_jobs
scores                  reports                notifications
```

See `TASKS.md` for the full column definitions of each table.

---

## Background Jobs (Queue)

**No Redis or Supervisor** — both require sudo and are unavailable on the shared dev server. All async work uses a **DB-based queue** instead.

### General Queue System

1. Any service calls `QueueService::push(type, payload)` → inserts a row into `queue_jobs`
2. A cron job runs `php spark queue:work` every minute (user-level crontab, no sudo needed)
3. `QueueWork` command claims the next `pending` row with `SELECT … FOR UPDATE`, runs the job, marks it `done` or `failed` (retries up to 3×)

### Automation Queue System (SEO Automation)

The SEO automation system has a dedicated queue separate from the general job queue:

**Phase 1 (Current):** Page-load triggered processing
- Triggered when users visit the SEO Automation dashboard
- Processes pending automation jobs once per day per website
- Uses localStorage to track last trigger time
- Calls `POST /api/v1/automation/process` endpoint with website ID

**Phase 2 (Planned):** Cron-based processing
- Add cron job to run automation processor daily/hourly
- Reliable scheduled processing independent of page visits
- See [AUTOMATION_PHASE_2_CRON_SETUP.md](AUTOMATION_PHASE_2_CRON_SETUP.md) for setup instructions

**Queue table:** `wc_automation_queue`
- Tracks audit crawls, opportunity generation, rank syncing jobs
- Statuses: pending → processing → completed (or failed)
- Auto-retries failed jobs up to 3 times
- Tracks last run time in `wc_websites.last_automation_run`

**Commands:**
```bash
# Process pending automation jobs
php spark automation:process --limit=10

# View pending jobs
mysql -h 127.0.0.1 -u dev -p'password' cms -e "SELECT * FROM wc_automation_queue WHERE status='pending';"

# Log output to file
php spark automation:process --limit=10 >> writable/logs/automation_queue.log 2>&1
```

### Setting up the crontab (General Queue)

```bash
crontab -e
```

Add:

```
* * * * * php /var/www/php82/<your-username>/webcrawlers_latest.com/spark queue:work >> writable/logs/worker.log 2>&1
0 2 * * * php /var/www/php82/<your-username>/webcrawlers_latest.com/spark rank:check
```

### Job classes

| Class | Trigger | Does |
|---|---|---|
| `CrawlSiteJob` | project creation / manual scan | Crawls domain, stores pages + errors |
| `GenerateScoresJob` | after crawl | Computes health score from crawl data |
| `GoogleSyncJob` | after GSC/GA4 connection | Pulls GSC / GA4 / GBP data |
| `KeywordJob` | on demand | Calls DataForSEO, stores results |
| `ReportJob` | scheduled | Assembles and stores report data |

All job classes live in `app/Jobs/`. The `spark` commands are in `app/Commands/`.

---

## Authentication

### Dashboard UI (session-based)

- Sign in → OTP → `session()->set(['user_id' => ..., 'logged_in' => true])`
- Redirect lands on `/dashboard`
- `AuthFilter` checks `session('logged_in')` on every `/dashboard/*` request; unauthenticated → redirect to `/signin`

### API (Bearer token)

- After login the `Dashboard` controller generates a **signed stateless token** per request:
  ```
  base64(JSON{uid, iat}) . "." . HMAC-SHA256(payload, WC_ENCRYPTION_KEY)
  ```
- The token is injected into every page as `window.WC.apiToken` (via `<meta name="api-token">`)
- Every `fetch()` call uses `window.wcFetch()` which sends `Authorization: Bearer <token>`
- `ApiAuthFilter` validates the HMAC and expiry (24 h); falls back to session auth when no Bearer header is present (for same-origin dashboard calls)
- `BaseApiController::resolveUserId()` decodes user ID from the token or session

This design means the same API works for:
- Dashboard JS (session or token)
- Future mobile apps (token only)
- Future public API (token only, separate `api_tokens` table to be added)

---

## Project Architecture

```
Public marketing site
        │
        ▼
  Pages / Auth controllers
        │
        ▼
Dashboard controller  ──────────────►  /dashboard/*  (HTML, session-protected)
        │                                     │
        │                                     ▼
        │                           Views render layout shell
        │                           JS calls wcFetch() on load
        │
        ▼
API v1 controllers  ────────────────►  /api/v1/*  (JSON, token-protected)
  ├─ WebsitesController
  ├─ AutomationController  ◄─── Triggers job queue processing
  └─ ... (other endpoints)
        │
        ▼
  Service Layer
  ┌─────────────────────────────────────────┐
  │  ProjectService   DashboardService      │
  │  CrawlService     KeywordService        │
  │  ContentService   AIService             │
  │  GoogleService    DataForSEOService     │
  │  QueueService     CRMService            │
  │  AuditSyncService                       │
  └─────────────────────────────────────────┘
        │
        ▼
     MySQL  ◄──►  DB Queues
        │         ├─ queue_jobs (general jobs)
        │         └─ wc_automation_queue (automation jobs)
        │
        ▼
  PHP CLI Workers (cron-triggered)
        │
        ├─► spark queue:work (general jobs)
        └─► spark automation:process (automation jobs)
        │
        ▼
  ┌─────────────────────────────────────────┐
  │  Google APIs  (GSC · GA4 · GBP)        │
  │  OpenAI       (GPT-4o)                 │
  │  DataForSEO   (Keywords · SERP · BL)   │
  │  Own Crawler  (Technical audit)        │
  └─────────────────────────────────────────┘
```

---

## Folder Structure

```
app/
  Commands/
    QueueWork.php            spark queue:work (general job processor)
    AutomationProcess.php    spark automation:process (SEO automation)
  Config/
    App.php           baseURL + environment detection
    Constants.php     API keys, WC_* constants, DB path
    Encryption.php    references WC_ENCRYPTION_KEY
    Filters.php       registers 'auth' and 'apiauth' filter aliases
    Routes.php        all routes — public, /dashboard/*, /api/v1/*
  Controllers/
    Auth.php          sign-up, sign-in, OTP, password reset
    Dashboard.php     serves all /dashboard/* HTML pages
    Pages.php         public marketing pages
    Lead.php          contact / book-demo form submissions
    Api/
      V1/
        BaseApiController.php       token decode, JSON helpers
        DashboardController.php     GET /api/v1/dashboard
        WebsitesController.php      website CRUD, automation toggle
        AutomationController.php    POST /api/v1/automation/process
  Database/
    Migrations/
      2026-08-20-000001_CreateWcAutomationQueueTable.php  automation queue
      2026-08-20-000002_AddLastAutomationRunToWebsites.php last run tracking
    Seeds/
  Filters/
    AuthFilter.php    session guard for /dashboard/*
    ApiAuthFilter.php Bearer token + session guard for /api/v1/*
  Jobs/               background job classes (CrawlSiteJob, etc.)
  Libraries/
    AuditSyncService.php  handles audit crawl queuing + DataForSEO integration
  Models/
    UserModel.php
    WebsiteModel.php
    AutomationQueueModel.php  job queue operations
    OtpCodeModel.php
    PasswordResetModel.php
  Services/           business logic layer (to be built per TASKS.md)
  Views/
    dashboard/
      layout/
        top.php       full HTML shell — sidebar, topbar, meta tags
        bottom.php    closing tags + wcFetch() bootstrap JS
        nav.php       sidebar nav map with site_url() hrefs
        components.php  dash_*() render helpers
      index.php       Dashboard page (fetches /api/v1/dashboard)
      stub.php        placeholder for pages not yet built
    signin.php  signup.php  (public auth views)

public/
  index.php           entry point — sets CI_ENVIRONMENT, boots CI4
  webcrawlers-dashboard-assets/
    assets/css/       dashboard.css + per-page CSS
    assets/js/
      seo-automation.js  SEO automation UI + auto-trigger logic
      dashboard.js    general dashboard JS
    *.php             original design reference files (not served by CI4)
    includes/         original layout includes (reference only)

Root Files:
  README.md           this file
  AUTOMATION_PHASE_2_CRON_SETUP.md  automation cron setup guide
  TASKS.md            project development checklist
  JIRA_TICKETS.md     ticket tracking
```

---

## Adding a New Dashboard Page

1. **Add the route** in `app/Config/Routes.php` inside the `dashboard` group:
   ```php
   $routes->get('my-page', 'Dashboard::myPage');
   ```

2. **Add the controller method** in `app/Controllers/Dashboard.php`:
   ```php
   public function myPage(): string
   {
       return $this->render('dashboard/my-page', [
           'active'     => 'my-page',   // matches key in nav.php
           'page_title' => 'My Page',
           'extra_css'  => ['my-page'], // loads assets/css/my-page.css
           'extra_js'   => ['my-page'], // loads assets/js/my-page.js
       ]);
   }
   ```

3. **Add the nav entry** in `app/Views/dashboard/layout/nav.php` if it needs a sidebar link.

4. **Create the view** at `app/Views/dashboard/my-page.php`.  
   Use the design reference at `public/webcrawlers-dashboard-assets/<page>.php` as the template — extract everything between `include 'includes/layout-top.php'` and `include 'includes/layout-bottom.php'`, replace `dashboard.php` hrefs with `site_url('dashboard/...')`, and use `dash_*()` helpers from `layout/components.php`.

   Until the view file exists, the page renders `stub.php` automatically (no crash).

5. **Add the API endpoint** if the page needs data — see below.

---

## Adding a New API Endpoint

1. Create a controller in `app/Controllers/Api/V1/` extending `BaseApiController`:
   ```php
   class MyController extends BaseApiController
   {
       public function index(): ResponseInterface
       {
           // $this->apiUserId is the authenticated user
           return $this->ok(['data' => []]);
       }
   }
   ```

2. Register the route in `app/Config/Routes.php` inside the `api/v1` group:
   ```php
   $routes->get('my-resource', 'MyController::index');
   ```

3. Call it from the view JS using `wcFetch()`:
   ```js
   wcFetch('my-resource').then(r => r.json()).then(data => { /* render */ });
   ```
   `wcFetch()` automatically attaches the Bearer token and CSRF header.

**Example: Automation Processing Endpoint**

The automation system provides `POST /api/v1/automation/process` to trigger job queue processing:

**Controller:** `app/Controllers/Api/V1/AutomationController.php`
```php
public function process(): ResponseInterface
{
    $body = $this->request->getJSON(true) ?? [];
    $websiteId = (int) ($body['website_id'] ?? 0);
    
    // Validate website, check rate-limiting (24-hour window)
    // Run automation queue processor
    // Update last_automation_run timestamp
    
    return $this->respond([
        'message'        => 'Automation queue processed successfully',
        'jobs_processed' => 2,
        'processed'      => true,
    ]);
}
```

**Route:** `app/Config/Routes.php`
```php
$routes->post('automation/process', 'AutomationController::process');
```

**Frontend call:** `public/webcrawlers-dashboard-assets/assets/js/seo-automation.js`
```js
wcFetch('automation/process', {
    method: 'POST',
    body: JSON.stringify({ website_id: p.id })
})
  .then(r => r.json())
  .then(data => console.log(data.message))
  .catch(err => console.warn(err));
```

---

## Design Assets

The UI design lives in `public/webcrawlers-dashboard-assets/`. These are **reference files only** — they are not CI4 views and are not served as part of the application. They exist so developers can view the intended design directly in a browser without logging in.

| Asset | Path |
|---|---|
| All page designs | `public/webcrawlers-dashboard-assets/*.php` |
| Dashboard CSS | `public/webcrawlers-dashboard-assets/assets/css/` |
| Dashboard JS | `public/webcrawlers-dashboard-assets/assets/js/` |
| Layout shell | `public/webcrawlers-dashboard-assets/includes/` |

When building a new view, open the matching `.php` design file side-by-side with your new CI4 view. The CI4 layout (`top.php` / `bottom.php`) already produces identical HTML to the design's `layout-top.php` / `layout-bottom.php`.

---

## External Integrations

### MVP Strategy: DataForSEO + OpenAI Only

**Important:** For MVP, only **DataForSEO** and **OpenAI** are required. Google APIs (GSC/GA4/GBP) are optional post-MVP enhancements.

**See:** [DATAFORSEO_VS_GOOGLE_APIS_ANALYSIS.md](DATAFORSEO_VS_GOOGLE_APIS_ANALYSIS.md) for detailed feature mapping and rationale.

### Integration Table

| Feature | Service | Constant | MVP? | Priority |
|---|---|---|---|---|
| Technical audit, crawl | DataForSEO On-Page API | `WC_DATAFORSEO_LOGIN` / `_PASSWORD` | ✅ YES | HIGH |
| Keyword research, gap, rank tracking | DataForSEO APIs | `WC_DATAFORSEO_LOGIN` / `_PASSWORD` | ✅ YES | HIGH |
| AI content generation, assistant | OpenAI GPT-4o | `WC_OPENAI_API_KEY` | ✅ YES | HIGH |
| Search Console data | Google GSC API (OAuth2) | `WC_GOOGLE_CLIENT_*` | ⚠️ OPTIONAL | LOW |
| Analytics traffic | Google GA4 API (OAuth2) | `WC_GOOGLE_CLIENT_*` | ⚠️ OPTIONAL | LOW |
| Local business data | Google Business Profile API (OAuth2) | `WC_GOOGLE_CLIENT_*` | ⚠️ OPTIONAL | LOW |
| CRM (leads, demos) | HubSpot or Freshdesk | `WC_HUBSPOT_API_KEY` or `WC_FRESHDESK_*` | ✅ YES | MEDIUM |

### Service Wrapper Classes (Priority Order)

**MVP (Required):**
- `DataForSEOService.php` — HTTP client for DataForSEO API (technical audit, keywords, rank tracking)
- `AIService.php` — OpenAI wrapper with retry (content generation)
- `CRMService.php` — push leads to HubSpot / Freshdesk

**Core Infrastructure:**
- `QueueService.php` — DB queue push / pop
- `CrawlEngine.php` — own PHP crawler for technical audit

**Post-MVP (Optional):**
- `GoogleService.php` — OAuth2 + GSC / GA4 / GBP (add after MVP validation)

---

## Development Phases

| Phase | Features | Status |
|---|---|---|
| **1** | Auth, Homepage, Project creation, Dashboard shell | In progress |
| **2** | Crawl engine, crawl history, crawl monitoring | Planned |
| **3** | ~~Google integrations~~ (Post-MVP enhancement) | Deferred |
| **4** | Keyword research, gap analysis, rank tracking | Planned |
| **5** | AI assistant, AI content writer, content briefs | Planned |
| **6** | Automation (queue processing), Reports, Notifications | In progress |

**Phase 6 - Automation (Current Work):**
- ✅ DB-based automation queue (`wc_automation_queue`)
- ✅ Page-load triggered processing (once per day per site)
- ✅ Resume Automation button → queues jobs → processes via API
- ✅ DataForSEO audit crawl integration
- 🚀 Phase 2: Add cron-based processing (see [AUTOMATION_PHASE_2_CRON_SETUP.md](AUTOMATION_PHASE_2_CRON_SETUP.md))

See `TASKS.md` for the full task breakdown per JIRA ticket.  
See `JIRA_TICKETS.md` for the original ticket requirements and acceptance criteria.


- PHP 8.2
- CodeIgniter 4
- MySQL
- DB-based Queue (general + automation queues)
- REST APIs
- Google APIs (Search Console, GA4, Business Profile)
- OpenAI API (GPT-4)
- DataForSEO API (Keywords, SERP, Backlinks, Audit Crawl)

---

# Architecture

```
           Frontend (HTML / Bootstrap)

                     │
              REST APIs (CodeIgniter 4)

                     │
           ┌─────────────────────┐
           │    Service Layer    │
           │  ProjectService     │
           │  CrawlService       │
           │  KeywordService     │
           │  ContentService     │
           │  AIService          │
           │  GoogleService      │
           │  QueueService       │
           │  AuditSyncService   │
           └─────────────────────┘

                     │
           ┌─────────────────────┐
           │  MySQL              │
           │  (queue_jobs table) │
           └─────────────────────┘

                     │
           ┌─────────────────────┐
           │  PHP CLI Workers    │
           │  (cron every min)   │
           │  CrawlSiteJob       │
           │  GenerateScoresJob  │
           │  GoogleSyncJob      │
           │  KeywordJob         │
           │  ReportJob          │
           └─────────────────────┘

                     │
           ┌─────────────────────────────────┐
           │  External Integrations          │
           │  Google APIs · OpenAI           │
           │  DataForSEO · Own Crawl Engine  │
           └─────────────────────────────────┘
```

---

# Recommended Folder Structure

```
application/

controllers/

models/

services/

repositories/

libraries/

workers/

jobs/

helpers/

config/
```

Example Services

```
ProjectService.php

DashboardService.php

CrawlService.php

KeywordService.php

GoogleService.php

ContentService.php

ReportService.php

AIService.php

QueueService.php
```

---

# Database Design

## Core Tables

```
users

projects

project_settings

project_integrations

crawl_runs

crawl_pages

crawl_changes

crawl_errors

keywords

keyword_history

keyword_rankings

reports

scores

notifications

queue_jobs

wc_websites            (SEO dashboard sites)

wc_automation_queue    (automation job queue)
```

---

# Queue System

**Approach: DB-based queue + user-level cron (no Redis, no Supervisor)**

**General Queue:**
```
queue_jobs table (MySQL)

└── status: pending | running | done | failed
└── attempts: int (max 3 retries)
└── run_at: datetime

QueueService::push(type, payload)  → inserts row
QueueService::pop(type)            → SELECT ... FOR UPDATE, claims row

spark queue:work  → cron every minute
spark rank:check  → cron daily at 02:00
```

**Automation Queue:**
```
wc_automation_queue table (MySQL)

└── website_id: int (FK to wc_websites)
└── user_id: int
└── job_type: enum (audit_crawl, opportunity_gen, rank_sync)
└── status: enum (pending, processing, completed, failed)
└── retry_count: int (max 3 retries)
└── last_error: longtext
└── created_at: datetime
└── processed_at: datetime

AutomationQueueModel::enqueue()      → inserts job
AutomationQueueModel::getPendingJobs() → queries pending jobs

spark automation:process            → cron job (Phase 2)
POST /api/v1/automation/process     → HTTP trigger (Phase 1)
```

Job Types

```
crawl_site

technical_audit

generate_scores

keyword_research

rank_tracking

audit_crawl           (automation - DataForSEO audit)

opportunity_gen       (automation - opportunity generation)

rank_sync             (automation - rank tracking sync)


sync_search_console

sync_ga4

sync_business_profile

generate_ai_content

verify_pixel

send_notifications
```

---

# REST APIs

```
GET     /dashboard

GET     /projects

POST    /projects

PUT     /projects/{id}

DELETE  /projects/{id}

GET     /crawl/history

GET     /crawl/diff

POST    /crawl/start

POST    /crawl/retry

GET     /keywords

POST    /keywords/search

GET     /reports

POST    /ai/chat
```

---

# Option A
# Using DataForSEO APIs

## Recommended Integrations

| Feature | API |
|----------|-----|
| Keyword Research | DataForSEO |
| Keyword Gap | DataForSEO |
| Search Volume | DataForSEO |
| SERP | DataForSEO |
| Rank Tracking | DataForSEO |
| Competitors | DataForSEO |
| Domain Metrics | DataForSEO |
| Backlinks | DataForSEO |
| Google Search Console | Google API |
| GA4 | Google API |
| Business Profile | Google API |
| AI Assistant | OpenAI |
| Content Writer | OpenAI |
| Crawl | Own Crawler |

---

## Feature Mapping

### Homepage

```
Dashboard API

↓

Project Summary

↓

Keyword Summary

↓

Reports

↓

Recent Crawls
```

---

### Project Creation

```
Validate Domain

↓

Save Project

↓

Initialize Crawl

↓

Queue Jobs
```

---

### Crawl

```
Start Crawl

↓

Store Crawl Pages

↓

Generate Issues

↓

Generate Scores
```

---

### Keyword Research

```
Frontend

↓

Keyword API

↓

DataForSEO

↓

Store Response

↓

Display
```

---

### Rank Tracking

```
Scheduled Job

↓

DataForSEO Rank API

↓

Save Rankings

↓

Generate Trends
```

---

### Keyword Gap

```
Our API

↓

DataForSEO

↓

Compare Domains

↓

Store Results
```

---

### AI Content

```
OpenAI

↓

Generate Content

↓

Store Draft

↓

User Approval
```

---

### Reports

```
Crawl

↓

Keyword Data

↓

Google APIs

↓

AI Summary

↓

Generate Report
```

---

# Advantages

- Faster development
- Reliable keyword data
- Accurate SERP
- Backlinks
- Competitor analysis
- Less maintenance

---

# Disadvantages

- Monthly API cost
- Vendor dependency
- API limits

---

# Option B
# Without DataForSEO

Everything built in-house.

---

## Integrations

| Feature | Source |
|----------|--------|
| Keyword Research | Google Autocomplete |
| Keyword Suggestions | Google Suggest |
| Search Console | Google API |
| Analytics | Google API |
| Business Profile | Google API |
| AI | OpenAI |
| Crawl | Own Crawler |
| SERP | Own Scraper |
| Rank Tracking | Own Scheduler |
| Backlinks | Not Available (or integrate later) |

---

# Crawl Engine

Recommended

```
PHP CLI Worker

or

Node Crawlee

or

Headless Chrome
```

Store

```
URL

Status

Title

Meta

Canonical

Schema

Links

Images

Heading Structure
```

---

# Keyword Research

Possible Sources

```
Google Suggest

People Also Ask

Related Searches

Search Console

User Search History
```

Generate

```
Keyword Ideas

Clusters

Long Tail Keywords
```

---

# Rank Tracking

```
Daily Cron

↓

Search Google

↓

Find Position

↓

Store

↓

Generate Trend
```

Need

```
Proxy Rotation

Captcha Handling

User Agents
```

---

# Keyword Gap

Own Logic

```
Project Keywords

VS

Competitor Keywords

↓

Difference

↓

Suggestions
```

Competitor keywords must come from

```
Search Console

Manual Import

Own Crawl
```

Accuracy will be lower.

---

# Content Briefs

Using

```
Top Ranking Pages

↓

Own Crawl

↓

Extract

Headings

Entities

FAQs

Schema

Word Count

↓

AI Summary
```

---

# Technical Audit

Own crawler generates

```
Broken Links

Duplicate Titles

Duplicate Meta

404

Redirect Chains

Missing H1

Missing Alt

Large Images

Canonical Issues

Schema Errors
```

---

# Crawl Monitoring

Store every crawl

```
Run

Pages

Errors

Coverage

Duration
```

---

# Crawl Diff

Compare

```
Run 10

VS

Run 11
```

Generate

```
Added URLs

Removed URLs

Modified URLs

Changed Metadata

Changed Content

Changed Schema
```

---

# Reports

Generate from

```
Crawl

↓

Technical Issues

↓

Google Data

↓

AI Summary
```

---

# Advantages

- No recurring SEO API costs
- Complete control
- Fully customizable
- Better long-term ownership

---

# Disadvantages

- Longer development time
- SERP scraping complexity
- Proxy management
- Lower keyword coverage
- No backlink database
- Harder to match Semrush/Ahrefs quality

---

# Recommended Hybrid Architecture ⭐⭐⭐⭐⭐

**Chosen: Option A (DataForSEO) + Own Crawl Engine, DB queue, no Redis/Supervisor**

```
Frontend (Bootstrap)

↓

CodeIgniter 4 APIs

↓

Service Layer

↓

MySQL  ←→  DB Queue (queue_jobs)

↓

PHP CLI Workers (cron-triggered)

↓

--------------------------------------------------

Google APIs (Search Console · GA4 · Business Profile)

OpenAI (GPT-4 — AI Assistant · Content Writer)

DataForSEO (Keywords · SERP · Rank Tracking · Backlinks)

Own Crawl Engine (Technical Audit · Crawl Diff · Scores)

--------------------------------------------------
```

---

# Suggested Development Phases

## Phase 1

- Authentication
- Homepage
- Project Management
- Dashboard

---

## Phase 2

- Crawl Engine
- Crawl History
- Crawl Monitoring

---

## Phase 3

- Google Integrations
- Installation Flow

---

## Phase 4

- Keyword Research
- Keyword Gap
- Rank Tracking

---

## Phase 5

- AI Assistant
- AI Content Writer
- Content Briefs

---

## Phase 6

- Reports
- Automation
- Notifications

---

# Final Recommendation

## MVP

- CodeIgniter 4 (PHP 8.2)
- MySQL (DB-based queue via `queue_jobs` table)
- PHP CLI Workers via user-level cron (no Supervisor, no Redis)
- Google APIs (Search Console, GA4, Business Profile)
- OpenAI (GPT-4)
- DataForSEO (Keywords, SERP, Rank Tracking, Backlinks)
- Own Crawl Engine (technical audit, crawl diff, scoring)
- Bootstrap (frontend)

This delivers the fastest path to a production-ready SEO platform while working within existing server constraints (no sudo, no Redis, no Supervisor).

## Long-Term

Gradually replace DataForSEO modules with proprietary services where it provides strategic advantage (e.g., keyword clustering, SERP analysis), while continuing to rely on Google APIs and the own crawler for first-party data and technical audits.