# crawl — repo context

> Curated overview for code_search. Source: `direqt-search/crawl` (GitHub). Node.js/TypeScript.

## Purpose
**Web crawler** — maintains local representations of target websites. Provides
reliable access to page **metadata**, **content**, and the **links** between
pages. Clients create **Site** resources (configure the target site + how to
crawl it); the crawler exposes **Page** resources for querying page data and
change notifications.

Crawled Pages are **not** end-user-searchable directly — they are the technical
substrate that higher-level components (index, search) turn into searchable
indexes and derivative content.

## Structure (top-level dirs)
- `crawl-api/` — the Crawl API surface (Site / Page resources)
- `src/` — crawler implementation (fetching, parsing, link extraction)
- `config/`, `scripts/`, `test/`

## Tech
Node.js 22 · TypeScript · Express · MongoDB (Mongoose) · `cheerio` (HTML parsing) ·
`@rowanmanning/feed-parser` (feeds) · `fasttext.wasm` (language detection) ·
GCP (Pub/Sub, Tasks, Storage, Secret Manager). Depends on `@direqt-internal/{iam-api,service-framework}`.
