# content-id — repo context

> Curated overview for code_search. Source: `publicgoodsw/content-id` (GitHub). Python service.

## Purpose
**Content ID Service (CID)** — content extraction, de-duplication, and storage of
articles + metadata for media-partner content. A core component for identifying,
analyzing, and categorizing web content (contextual advertising + search.com widget).

## Processing pipeline
1. **Content extraction/parsing** — articles from URLs, with per-publisher handling
2. **De-duplication** — simhash-based matching to skip duplicate content
3. **ML classification** — multiple models (IAB taxonomy, UN SDGs, PGN taxonomy, purpose-driven topics)
4. **LLM enrichment** — GPT-powered keywords, categorization, example questions (via the `llm-content-enrichment` Lambda)
5. **Rules engine** — content filtering/categorization
6. **Related content** — hybrid model + keyword "read more" recommendations (search.com widget)
7. **Bypass mode** — analyze external content without partner restrictions (ADX categories for admedia)

## Structure
- `contentid/` — the service implementation
- `migrations/` — DB migrations · `opt/` · `cron.yaml`
- `Dockerfile`, `docker-compose.yml`, `Pipfile` — containerized Python app
- `llm-backfill.py` — LLM backfill script

## Tech
Python · Pipenv · Docker · MySQL/DB (migrations) · simhash · ML models · OpenAI GPT.
Config for LLM per partner lives in the DB `llm_config` table.
