|
Download README.md from cx-cmu/ReScraper: direct link, hf CLI and curl.
- Browser
- Download file 2.66 kB
-
https://huggingface.co/cx-cmu/ReScraper/resolve/main/README.md
- Command line
-
hf download hf://cx-cmu/ReScraper/README.md
-
curl -L -o README.md https://huggingface.co/cx-cmu/ReScraper/resolve/main/README.md
2.66 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| tags: | |
| - pretraining-data | |
| - data-curation | |
| - web-scraping | |
| base_model: | |
| - Qwen/Qwen3-0.6B | |
| # ReScraper | |
| ReScraper is a 0.6B refiner that replaces the heuristic HTML-to-text stack of a pretraining | |
| data pipeline — a rule-based scraper followed by rule-based cleaning filters — with a single | |
| small language model. | |
| - **Data:** [`cx-cmu/ReScraper-Data`](https://huggingface.co/datasets/cx-cmu/ReScraper-Data) | |
| - **Code:** [`cxcscmu/ReScraper`](https://github.com/cxcscmu/ReScraper) | |
| ## Contents | |
| - [What the model does](#what-the-model-does) | |
| - [Files](#files) | |
| - [Usage](#usage) | |
| - [Training](#training) | |
| - [License](#license) | |
| ## What the model does | |
| The model reads the visible text of a raw web page, rendered one block per line with line ids | |
| `<lid:n>`, and generates a short program: | |
| 1. An `<extract>` block of line removals (`rm 1-4`, `rm 6`, …) that strips headers, menus, | |
| boilerplate and other non-content lines. | |
| 2. One operation tag for the page: | |
| | Tag | Meaning | | |
| |---|---| | |
| | `<keep>` | keep the extracted text as it is | | |
| | `<edit>` | remove further noisy lines and spans inside the page | | |
| | `<delete>` | drop the page entirely | | |
| | `<rewrite>` | rewrite a poorly written but informative page | | |
| An executor applies the program to the page, so extraction and cleaning happen in one model | |
| instead of a cascade of separate stages. | |
| ## Files | |
| ``` | |
| models/ReScraper/ the refiner (Stage 2 checkpoint), in Hugging Face format | |
| ``` | |
| ## Usage | |
| ```python | |
| from vllm import LLM, SamplingParams | |
| llm = LLM(model="cx-cmu/ReScraper", dtype="bfloat16", max_model_len=32768) | |
| params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=3072) | |
| system = open("prompts/student_system_stage2.txt").read() # from the code repository | |
| page = "<lid:1> ...\n<lid:2> ...\n" # rendered page, one block per line | |
| out = llm.chat( | |
| [[{"role": "system", "content": system}, {"role": "user", "content": page}]], | |
| params, | |
| ) | |
| print(out[0].outputs[0].text) | |
| ``` | |
| The system prompt and the executor that applies the generated program are in the code | |
| repository, together with the full pipeline (rendering, inference, post-filtering, | |
| deduplication and tokenization). | |
| ## Training | |
| The model is trained in two supervised stages, distilling three teachers: a main-content | |
| extractor, a refinement teacher that decides the operation and edits the text, and a rewriting | |
| teacher for pages that are informative but poorly written. Both training sets are released in | |
| [`cx-cmu/ReScraper-Data`](https://huggingface.co/datasets/cx-cmu/ReScraper-Data). | |
| ## License | |
| Apache 2.0. |