# RAG Service

Track 2 orchestrator for the Adpilot retrieval layer: public `/search`, `/query`, and `/repos` APIs, Tier-1 intent analysis, and delegation to the embedding + retrieval service for hydrated context.

The RAG service does **not** re-hydrate chunk text from Mongo. It calls `POST /internal/retrieve` on the embedding engine (`:6004`) and adds intent metadata, LLM context assembly (P4), and answer synthesis.

## Configuration

Copy `.env.example` to `.env`. Key groups:

| Group | Variables |
| --- | --- |
| App | `APP_NAME`, `ENVIRONMENT` (or `ENV`), `PORT`, `LOG_LEVEL` |
| MongoDB | `MONGO_URI`, `DATABASE_NAME`, `INDEXING_RUNS_DATABASE` |
| Retrieval | `EMBEDDING_RETRIEVAL_SERVICE_URL`, `EMBEDDING_RETRIEVAL_TIMEOUT_SEC`, `RETRIEVAL_CLIENT_MOCK`, `RETRIEVAL_MOCK_SCENARIO` |
| Repo resolution | `DEFAULT_REPO_ID`, `REPO_RESOLVER_LLM_ENABLED`, `REPO_RESOLVER_CONFIDENCE_THRESHOLD` |
| Audit | `QUERY_LOGS_ENABLED` |
| Intent / LLM | `INTENT_CLASSIFIER`, `OPENAI_API_KEY`, `LLM_MODEL`, `QUERY_DEFAULT_TOP_K`, `QUERY_DEFAULT_SCORE_THRESHOLD`, `GENERAL_LLM_FALLBACK_ENABLED` |

In `development` / `local`, missing secrets are allowed. In `production`, startup validation fails fast if required values are empty.

## Project Structure

- `app/core/config.py`: validated settings
- `app/intent/`: Tier-1 rule-based intent analyzer
- `app/retrieval/retrieval_client.py`: HTTP client to embedding-engine `/internal/retrieve`
- `app/repositories/repo_repository.py`: read-only repo list from `adpilot_repo_sync`
- `app/services/repo_resolver.py`: optional `repo_id` resolution + 409 clarification
- `app/api/routes/`: public search, query, repos endpoints

## Repo resolution

`repo_id` is **optional** on `/search` and `/query`. When omitted, the service resolves it from indexed repos (single-repo auto-pick, slug match in question, `DEFAULT_REPO_ID`, or LLM repo intent). If ambiguous with multiple repos, returns **409**:

```bash
curl -X POST localhost:6005/api/v1/search \
  -H 'Content-Type: application/json' \
  -d '{"question":"Where is ProcessDelta defined?"}'

# Turn 2 after user picks a repo:
curl -X POST localhost:6005/api/v1/search \
  -H 'Content-Type: application/json' \
  -d '{"repo_id":"ad/your-repo.com","question":"Where is ProcessDelta defined?"}'
```

## Requirements

```bash
pip install -r requirements.txt
pip install -r tests/requirements-test.txt
```

## Run Locally

```bash
uvicorn app.main:app --reload --port 6005
```

Check health:

```bash
curl http://localhost:6005/healthz
curl http://localhost:6005/readyz
curl http://localhost:6005/api/v1/version
```

Run unit tests:

```bash
pytest tests/unit -q
```

Run integration tests (mocked retrieval + LLM, no OpenAI or Mongo required):

```bash
pytest tests/integration -q
```

E2E against live embedding-engine (requires stack on `:6004` + indexed data):

```bash
# Ensure embedding-engine is up and RETRIEVAL_CLIENT_MOCK=false in .env
uvicorn app.main:app --port 6005

curl -X POST localhost:6005/api/v1/search \
  -H 'Content-Type: application/json' \
  -d '{"repo_id":"<your-repo-id>","question":"How does commit analysis get triggered?","snapshot_id":"<snapshot-id>"}'

curl -X POST localhost:6005/api/v1/query \
  -H 'Content-Type: application/json' \
  -d '{"repo_id":"<your-repo-id>","question":"Explain the indexing pipeline","snapshot_id":"<snapshot-id>"}'

curl -N -X POST localhost:6005/api/v1/query \
  -H 'Content-Type: application/json' \
  -d '{"repo_id":"<your-repo-id>","question":"Explain the indexing pipeline","options":{"stream":true},"snapshot_id":"<snapshot-id>"}'
```

For offline dev without embedding-engine, set `RETRIEVAL_CLIENT_MOCK=true` in `.env`.

## Docker

Standalone:

```bash
cp .env.example .env
mkdir -p logs
docker build -t adpilot-rag-service .
docker run --env-file .env -p 6005:6005 adpilot-rag-service
```

Via composer stack (see `adpilot-common-composer.com/docker-compose.yml`):

```bash
cp services/adpilot-rag-service.com/.env.example services/adpilot-rag-service.com/.env
# Set OPENAI_API_KEY in the service .env for /query
docker compose up -d --build rag-service
docker compose logs -f rag-service
```

## API Endpoints (D1)

| Method | Path | Description |
| --- | --- | --- |
| `GET` | `/healthz` | Liveness probe |
| `GET` | `/readyz` | Readiness (Mongo + embedding-engine) |
| `GET` | `/api/v1/version` | Service metadata |
| `POST` | `/api/v1/search` | Retrieval-only search with intent |
| `POST` | `/api/v1/query` | Full RAG (non-streaming + SSE via `options.stream=true`); falls back to general LLM when retrieval is empty unless `GENERAL_LLM_FALLBACK_ENABLED=false` |
| `GET` | `/api/v1/repos` | Indexed repositories |

## Related Docs

- [Retrieval Layer HLD](../../docs/retrieval-layer-hld.md) — §4.2 RAG orchestrator, §7 public API contracts
- [Embedding + Retrieval Contracts](../../docs/retrieval-embedding-service-contracts.md) — `/internal/retrieve` schema

## Dependencies

- **Embedding engine** (`adpilot-indexing-embedding-engine.com`, port `6004`) for vector search and Mongo hydration
- **MongoDB** — reads `repositories` / `indexing_runs` from `adpilot_repo_sync`; writes `query_logs` to `adpilot_rag` (P4/D5)
