# AI Assistant — Backend Service

## What is this?

This is the backend service for the VS Code AI Assistant extension. Its sole job is to act as a secure bridge between the VS Code extension and the OpenAI API.

The extension never calls OpenAI directly. It calls this backend. The backend validates the request, builds a structured prompt, calls ChatGPT, and returns a clean response.

---

## What problem does it solve?

Without a backend:
- The OpenAI API key would live inside the VS Code extension — visible to anyone who installs it
- Every user would need their own API key
- There would be no rate limiting, no logging, no control over how prompts are constructed

With this backend:
- The API key lives only on the server in `.env`
- All users share one controlled entry point
- Prompt engineering, validation, and error handling are centralised

---

## Architecture

```
VS Code Extension
      │
      │  POST /api/v1/ask  (or /ask/stream)
      ▼
┌─────────────────────┐
│   Fastify Backend   │
│                     │
│  1. Validate input  │
│  2. Build prompt    │
│  3. Call OpenAI     │
│  4. Parse response  │
│  5. Return result   │
└─────────────────────┘
      │
      │  chat.completions.create()
      ▼
   OpenAI API (ChatGPT)
```

---

## Tech Stack

| Technology | Purpose |
|---|---|
| **Fastify** | HTTP server framework |
| **TypeScript** | Type safety across the entire codebase |
| **Zod** | Runtime request validation with field-level errors |
| **OpenAI SDK** | Communicates with ChatGPT API |
| **Pino** | Structured logging (pretty in dev, JSON in prod) |
| **dotenv** | Environment variable management |

---

## Project Structure

```
src/
├── config/
│   ├── env.ts          — Loads and validates all environment variables at startup
│   └── logger.ts       — Pino logger config (pretty in dev, JSON in prod)
│
├── plugins/
│   └── security.ts     — CORS, Helmet (security headers), rate limiting
│
├── routes/
│   ├── health.ts       — GET /health — liveness probe
│   ├── ask.ts          — POST /api/v1/ask — full response
│   └── ask.stream.ts   — POST /api/v1/ask/stream — SSE streaming
│
├── schemas/
│   └── ask.schema.ts   — Single source of truth for all request/response types
│                          Contains Zod schemas, JSON Schema for AJV, and inferred TS types
│
├── services/
│   ├── llm.client.ts   — OpenAI client singleton
│   └── ask.service.ts  — Core business logic: prompt building, OpenAI calls, response parsing
│
├── types/
│   └── index.ts        — Shared TypeScript types
│
├── app.ts              — Fastify app factory (registers plugins and routes)
└── server.ts           — Entry point — starts the server, handles graceful shutdown
```

---

## How a Request Flows Through the System

```
1. Request arrives at POST /api/v1/ask

2. Fastify AJV (JSON Schema)
   └── Checks structure: required fields, correct types, max lengths
   └── Rejects instantly if malformed — no further processing

3. Zod validation
   └── Runs deeper rules: min/max length, enum values, trim whitespace
   └── Returns field-level errors: which field failed and why

4. ask.service.ts — buildPrompt()
   └── Combines type + query + context into a structured prompt string
   └── Only includes context fields that are non-empty

5. ask.service.ts — OpenAI call
   └── Sends prompt to gpt-4o-mini (configurable via .env)
   └── Requests JSON response format

6. ask.service.ts — parseCompletion()
   └── Strips any markdown wrapping ChatGPT may have added
   └── Parses JSON into { plan, code, references }
   └── Falls back gracefully if JSON is malformed

7. Route handler returns HTTP 200 with structured response
```

---

## Endpoints

### GET /health

Liveness probe. Returns immediately without calling OpenAI.

**Response**
```json
{
  "status": "ok",
  "timestamp": "2026-05-07T06:32:00.000Z",
  "uptime": 3600,
  "environment": "development"
}
```

---

### POST /api/v1/ask

Sends a prompt to ChatGPT and waits for the complete response before returning.

**Request**
```json
{
  "type": "explain",
  "query": "Why is this function returning undefined?",
  "context": {
    "filePath": "src/utils/parser.ts",
    "code": "function parse(input) { return input.value }",
    "selection": "input.value"
  },
  "history": [
    { "role": "user",      "content": "What does this file do?" },
    { "role": "assistant", "content": "This file exports a parse utility that..." }
  ]
}
```

**Response — 200**
```json
{
  "plan": "Step-by-step explanation...",
  "code": "fixed code or empty string",
  "references": ["https://..."]
}
```

---

### POST /api/v1/ask/stream

Same as `/ask` but streams tokens using Server-Sent Events (SSE) as ChatGPT generates them. Recommended for the VS Code extension UI.

**Request** — identical to `/ask`

**Response** — stream of SSE lines

```
data: {"type":"chunk","content":"The "}
data: {"type":"chunk","content":"function "}
data: {"type":"chunk","content":"fails because..."}
data: {"type":"done","data":{"plan":"...","code":"...","references":[]}}
```

Three possible event types:

| Event | When | Contains |
|---|---|---|
| `chunk` | Every token from OpenAI | `content` — text to append |
| `done` | Once, at the end | `data` — full `{ plan, code, references }` |
| `error` | On OpenAI failure | `message` — human-readable error |

---

## Request Field Rules

| Field | Required | Limit |
|---|---|---|
| `type` | YES | Must be `"explain"`, `"debug"`, or `"implement"` |
| `query` | YES | 1 – 4,000 characters |
| `context.filePath` | no | Max 500 characters |
| `context.code` | no | Max 50,000 characters |
| `context.selection` | no | Max 10,000 characters |
| `history` | no | Max 20 items; each item's `content` ≤ 10,000 characters |

Each `history` item must have:
- `role`: `"user"` or `"assistant"`
- `content`: the message text as shown in the chat UI

When provided, history is injected into the OpenAI message chain before the current user message, giving the model full context of the prior conversation turns.

Requests exceeding these limits are rejected with HTTP 400 before any OpenAI call is made.

---

## Error Responses

All errors follow the same shape:

```json
{
  "statusCode": 400,
  "error": "Validation Error",
  "message": "One or more fields are invalid.",
  "fields": [
    { "field": "query", "message": "query must not be empty" }
  ]
}
```

| Status | Meaning |
|---|---|
| `400` | Bad request — `fields` array shows exactly which field failed |
| `429` | Rate limited — 60 requests per minute per IP |
| `502` | OpenAI error — key invalid, rate limit, or OpenAI is down |
| `500` | Unexpected server error |

---

## How the Prompt is Built

The service constructs a prompt string from the request and sends it to ChatGPT. Example of what ChatGPT receives:

```
You are a senior software engineer helping a developer inside VS Code.

Task type: DEBUG

User request:
Why is this function returning undefined?

Context:
File: src/utils/parser.ts

```
function parse(input) { return input.value }
```

Selected code:
input.value

Respond ONLY with a valid JSON object in this exact shape...
```

ChatGPT is instructed to always reply in a specific JSON format. The backend parses that JSON and returns it as the API response.

---

## Environment Variables

All variables are validated at startup. The server refuses to start if any required variable is missing.

```env
NODE_ENV=development          # development | production | test
HOST=0.0.0.0
PORT=3000

OPENAI_API_KEY=sk-...         # Required — your OpenAI key
OPENAI_MODEL=gpt-4o-mini      # Model to use
OPENAI_MAX_TOKENS=2048
OPENAI_TEMPERATURE=0.3

CORS_ORIGIN=*                 # Comma-separated origins or * for all
RATE_LIMIT_MAX=60             # Requests per window per IP
RATE_LIMIT_WINDOW_MS=60000    # Window in milliseconds

LOG_LEVEL=info
```

---

## Running Locally

```bash
# Install dependencies
npm install

# Start development server with hot reload
npm run dev

# Type check
npm run type-check

# Build for production
npm run build

# Start production build
npm start
```

Server starts at `http://localhost:3000`.

---

## Security

| Measure | What it does |
|---|---|
| **Helmet** | Sets HTTP security headers on every response |
| **CORS** | Restricts which origins can call the API |
| **Rate limiting** | Blocks IPs that exceed 60 requests per minute |
| **Input validation** | All inputs validated and size-capped before processing |
| **API key** | Stored only in `.env` on the server — never exposed to clients |
| **Error messages** | 500 errors never leak internal details to the caller |

---

## What This Backend Does NOT Do

- It does not read files from disk — file content is sent by the extension
- It does not persist conversation history — history must be sent by the client on each request
- It does not authenticate users
- It does not modify code in the editor
- It does not understand the full codebase — only what is sent in the request
