ReScraper / README.md
yuzc19's picture
Update README.md
a15cfd6 verified
|
Raw History Blame Contribute Delete
2.66 kB
---
license: apache-2.0
language:
- en
pipeline_tag: text-generation
tags:
- pretraining-data
- data-curation
- web-scraping
base_model:
- Qwen/Qwen3-0.6B
---
# ReScraper
ReScraper is a 0.6B refiner that replaces the heuristic HTML-to-text stack of a pretraining
data pipeline — a rule-based scraper followed by rule-based cleaning filters — with a single
small language model.
- **Data:** [`cx-cmu/ReScraper-Data`](https://huggingface.co/datasets/cx-cmu/ReScraper-Data)
- **Code:** [`cxcscmu/ReScraper`](https://github.com/cxcscmu/ReScraper)
## Contents
- [What the model does](#what-the-model-does)
- [Files](#files)
- [Usage](#usage)
- [Training](#training)
- [License](#license)
## What the model does
The model reads the visible text of a raw web page, rendered one block per line with line ids
`<lid:n>`, and generates a short program:
1. An `<extract>` block of line removals (`rm 1-4`, `rm 6`, …) that strips headers, menus,
boilerplate and other non-content lines.
2. One operation tag for the page:
| Tag | Meaning |
|---|---|
| `<keep>` | keep the extracted text as it is |
| `<edit>` | remove further noisy lines and spans inside the page |
| `<delete>` | drop the page entirely |
| `<rewrite>` | rewrite a poorly written but informative page |
An executor applies the program to the page, so extraction and cleaning happen in one model
instead of a cascade of separate stages.
## Files
```
models/ReScraper/ the refiner (Stage 2 checkpoint), in Hugging Face format
```
## Usage
```python
from vllm import LLM, SamplingParams
llm = LLM(model="cx-cmu/ReScraper", dtype="bfloat16", max_model_len=32768)
params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=3072)
system = open("prompts/student_system_stage2.txt").read() # from the code repository
page = "<lid:1> ...\n<lid:2> ...\n" # rendered page, one block per line
out = llm.chat(
[[{"role": "system", "content": system}, {"role": "user", "content": page}]],
params,
)
print(out[0].outputs[0].text)
```
The system prompt and the executor that applies the generated program are in the code
repository, together with the full pipeline (rendering, inference, post-filtering,
deduplication and tokenization).
## Training
The model is trained in two supervised stages, distilling three teachers: a main-content
extractor, a refinement teacher that decides the operation and edits the text, and a rewriting
teacher for pages that are informative but poorly written. Both training sets are released in
[`cx-cmu/ReScraper-Data`](https://huggingface.co/datasets/cx-cmu/ReScraper-Data).
## License
Apache 2.0.