URL:     https://linuxfr.org/users/killruana/journaux/selenium-et-pandoc-pour-lire-des-nouvelles-hors-lignes
Title:   selenium et pandoc pour lire des nouvelles hors lignes
Authors: jtremesay
Date:    2026-10-08T16:31:12+02:00
License: CC By-SA
Tags:    chine, web-nouvelle, epub, selenium, pandoc et python
Score:   10


intro tl;dr: j'ai découvert le [c-drama](https://en.wikipedia.org/wiki/Chinese_television_drama) [Against the Current](https://en.wikipedia.org/wiki/Against_the_Current_(TV_series)). J'ai découvert qu'à la base c'est une [web-nouvelle](https://en.wikipedia.org/wiki/Web_fiction) publié en chinois. J'ai trouvé des [fan-lations](https://en.wikipedia.org/wiki/Fan_translation) en ligne dont celle [là](https://www.readthedrama.com/novels/against-the-current), mais personne n'ayant eu la sympathie de condenser ça en .epub pour lire facilement sur la liseuse.

Du coup, aujourd'hui, on va voir comment automatiser tout ça en moins de 100 lignes de python.

Il vous faut :

- python3
- [Selenium webdriver](https://www.selenium.dev/documentation/webdriver/), une lib pour piloter un navigateur web
- [Pandoc](https://pandoc.org/), le convertisseur universel pour générer l'epub à partir de nos chapitres markdown
- [tqdm](https://tqdm.github.io/), pour avoir une jolie barre de progression. Totalement optionnel

```python
from pathlib import Path
from subprocess import run

from selenium.webdriver import Firefox
from selenium.webdriver.common.by import By

try:
    from tqdm import tqdm
except ImportError:
    tqdm = lambda x: x

BASE_URL = "https://www.readthedrama.com"
NOVEL_SLUG = "against-the-current"
CHAPTERS = 358
OUTPUT_DIR = Path.cwd() / "output"
NOVEL_NAME = "Against the Current"
NOVEL_AUTHOR = "He Yan Shan"


def dl_chapter(driver: Firefox, chapter_num: int, output_dir: Path) -> None:
    chapter_path = output_dir / f"chapter_{chapter_num:03}.md"
    if chapter_path.exists():
        return

    driver.get(f"{BASE_URL}/novels/{NOVEL_SLUG}/chapters/{chapter_num}")
    driver.implicitly_wait(1)

    elements = driver.find_elements(By.CSS_SELECTOR, "article > p")
    novel_text = f"## Chapter {chapter_num}\n\n" + "\n".join(
        [element.text + "\n" for element in elements]
    )
    with open(chapter_path, "w", encoding="utf-8") as f:
        f.write(novel_text)


def dl_chapters(driver: Firefox, output_dir: Path) -> None:
    for i in tqdm(range(1, CHAPTERS + 1)):
        dl_chapter(driver, i, output_dir)


def generate_epub(output_dir: Path) -> None:
    with (output_dir / "metadata.yaml").open("w") as f:
        f.write(f"""\
---
title: "{NOVEL_NAME}"
author: "{NOVEL_AUTHOR}"
toc: true
toc-title: "Table of contents"
...
""")
    args = [
        "pandoc",
        "metadata.yaml",
        "-o",
        f"{NOVEL_SLUG}.epub",
        "--toc",
        "--split-level=1",
        "--file-scope",
    ] + [f"chapter_{i:03}.md" for i in range(1, CHAPTERS + 1)]
    run(args, cwd=output_dir, check=True)


def main() -> None:
    driver = Firefox()
    driver.install_addon("consent_o_matic-1.1.5.xpi")
    driver.install_addon("uBlock0_1.75.0.firefox.signed.xpi")

    output_dir = OUTPUT_DIR / NOVEL_SLUG
    output_dir.mkdir(parents=True, exist_ok=True)

    dl_chapters(driver, output_dir)
    driver.quit()

    generate_epub(output_dir)


if __name__ == "__main__":
    main()
```
