# SEO Automation System - Phase 2: Cron Setup Documentation

## Current Implementation (Phase 1)

The automation system processes jobs once per day on page load, using localStorage to track the last trigger time per website.

**How it works:**
- When the SEO Automation page loads, it checks localStorage for today's date
- If not found, it calls `/api/v1/automation/process` endpoint with the website ID
- The endpoint checks if 24 hours have passed since last run
- If yes, it runs the queue processor and updates `wc_websites.last_automation_run`
- If no, it returns a message indicating when the next run is allowed

**Files involved:**
- `public/webcrawlers-dashboard-assets/assets/js/seo-automation.js` - Triggers endpoint on page load
- `app/Controllers/Api/V1/AutomationController.php` - Handles automation processing requests
- `app/Database/Migrations/2026-08-20-000002_AddLastAutomationRunToWebsites.php` - Tracks last run time

---

## Phase 2: Add Cron Job for Reliable Daily Processing

When you're ready to move to Phase 2, add this to your user crontab:

```bash
crontab -e
```

Add this line to run automation processing daily at 2 AM:

```
0 2 * * * /usr/bin/php8.2 /var/www/php82/koushik/webcrawlers_latest.com/spark automation:process --limit=100 >> /var/www/php82/koushik/webcrawlers_latest.com/writable/logs/automation_queue.log 2>&1
```

**Or for every 6 hours:**

```
0 */6 * * * /usr/bin/php8.2 /var/www/php82/koushik/webcrawlers_latest.com/spark automation:process --limit=100 >> /var/www/php82/koushik/webcrawlers_latest.com/writable/logs/automation_queue.log 2>&1
```

**Or for every 5 minutes (frequent processing):**

```
*/5 * * * * /usr/bin/php8.2 /var/www/php82/koushik/webcrawlers_latest.com/spark automation:process --limit=100 >> /var/www/php82/koushik/webcrawlers_latest.com/writable/logs/automation_queue.log 2>&1
```

---

## Using HTTP Endpoint Instead of CLI

If you want external services to trigger automation (like AWS EventBridge, Google Cloud Scheduler, etc.):

```bash
# Call this endpoint daily or as often as needed
curl -X POST "http://your-domain.com/api/v1/automation/process" \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer YOUR_API_TOKEN" \
  -d '{"website_id": 12}'
```

Response example:
```json
{
  "message": "Automation queue processed successfully",
  "jobs_processed": 2,
  "processed": true
}
```

---

## Queue Processing Command Reference

### Process all pending jobs:
```bash
cd /var/www/php82/koushik/webcrawlers_latest.com
/usr/bin/php8.2 spark automation:process --limit=10
```

### Process with custom limit:
```bash
/usr/bin/php8.2 spark automation:process --limit=50
```

### Filter by job type (when implemented):
```bash
/usr/bin/php8.2 spark automation:process --limit=10 --job-type=audit_crawl
```

### Log to file:
```bash
/usr/bin/php8.2 spark automation:process --limit=10 >> writable/logs/automation_queue.log 2>&1
```

---

## Database Schema

**wc_automation_queue**
- `id` - Auto-increment job ID
- `website_id` - Foreign key to wc_websites
- `user_id` - User who initiated the job
- `job_type` - 'audit_crawl', 'opportunity_gen', or 'rank_sync'
- `status` - 'pending', 'processing', 'completed', or 'failed'
- `retry_count` - Number of retry attempts
- `last_error` - Error message if job failed
- `created_at` - When job was queued
- `processed_at` - When job completed/failed

**wc_websites changes**
- `automation_enabled` - Enable/disable automation for a site
- `last_automation_run` - Timestamp of last queue processing (Phase 1 tracking)

---

## Transition Plan from Phase 1 to Phase 2

1. Keep Phase 1 running (page load triggers)
2. Add cron job as fallback
3. Remove localStorage tracking from JavaScript
4. Rely solely on `wc_websites.last_automation_run` for rate limiting
5. Monitor logs for success

---

## Job Status Flow

```
pending → processing → completed
       → failed (with retry_count)
```

The processor automatically:
- Retries failed jobs up to 3 times
- Logs all errors
- Marks jobs as completed
- Updates last_automation_run timestamp after successful run
