# AGENTS.md — AI-facing documentation for this repo

(README.md is 100% human; keep it that way. Put all agent-written docs here.)

## What this project is

A dependency-free (stdlib-only) Python manga archiver + static site generator.
It downloads all chapters of a manga from a supported source into
`<archive_dir>/chapter_XXX/page_001.ext, ...`, and can render the archive as
a static HTML reading site.

Current archives on disk:
- `disastrous-necromancer/` — from drakecomic.org (Drake Scans).
- `return-of-the-runebound-professor/` — from mangaread.org.

## Layout

- `manga/core.py` — shared engine: polite fetching, resumable self-healing
  chapter downloads, `download_state.json` state file, `run(source, url)`.
- `manga/drake.py` — Drake Scans adapter. Chapter pages embed
  `ts_reader.run({"sources":[{"images":[...]}]})`; that list is the canonical
  page set. (Some chapter posts include images from other manga; we
  deliberately keep exactly what the site's reader lists.)
- `manga/mangaread.py` — MangaRead adapter (WordPress "Madara" theme):
  chapter list = `<li class="wp-manga-chapter">` items (newest first, so we
  reverse to reading order); pages = `wp-manga-chapter-img` img tags
  (`src` or `data-src`) inside the `reading-content` container, in display
  order. Chapter slugs may carry fix suffixes (`chapter-71-1`).
- `download.py` — CLI dispatcher over the adapters.
- `generate_html.py` — static reader pages; idempotent, works with any
  archive dir; single CLI arg (dir), default `disastrous-necromancer`.
- `TODO.md` — human notes.
- `download.log` — log of previous drakecomic downloads (gitignored).

## Usage

```
# Drake Scans (bare invocation = original default target, for backwards compat)
python3 download.py
python3 download.py drake [manga_url]

# MangaRead
python3 download.py mangaread https://www.mangaread.org/manga/<slug>/

# Static site for any archive dir
python3 generate_html.py [manga_dir]
```

Downloads are resumable: re-run any time; each chapter dir is reconciled with
the site's canonical page list (existing files matched by URL basename or
`page_NNN` name and renamed into place, missing pages fetched, stray image
files removed, progress recorded in `<archive_dir>/download_state.json`).
`.part` temp files are used so a crash mid-download never leaves a truncated
image at its final name.

## Politeness contract (important — do not weaken)

All network access goes through `manga/core.py:fetch`:
- single sequential connection, no concurrency;
- ~2s before each chapter-page fetch, ~0.75s before each image;
- longer breather (5s) every 25 images;
- exponential backoff on 5xx / network errors (up to 5 attempts);
- on HTTP 429: exponential backoff, and a hard exit (code 2) after 3
  consecutive 429s — the site asks us to stop, we stop.
- browser-like User-Agent and a Referer header on image requests (Madara
  hotlink protection expects the chapter page as referer).

When adding a new source, reuse `core.run` + the Source adapter interface
(`output_dir`, `chapter_dirname`, `get_chapters`, `get_pages`) instead of
rolling new fetch logic.

## Conventions

- Python 3.10+, stdlib only (`urllib`, `re`, `json`, `pathlib`). No
  third-party deps, no HTML parser library — the site markup is matched with
  narrow regexes scoped to the relevant containers, on purpose.
- Chapter dirs: `chapter_<n>`; mangaread fix chapters
  (`chapter-71-1` slugs) become `chapter_71_1` and are titled "Chapter 71.1".
  `generate_html.py` sorts chapter dirs numerically, handling the `_` suffix.
- Don't commit the image archives; `.gitignore` covers `*.part`, `download.log`,
  `nohup.out`, `__pycache__/`, and the archive dirs.
- Long runs: `nohup python3 download.py mangaread <url> >> download.log 2>&1 &`
  — the run is safe to interrupt and resume.
