--- license: apache-2.0 language: - en pipeline_tag: text-generation tags: - pretraining-data - data-curation - web-scraping base_model: - Qwen/Qwen3-0.6B --- # ReScraper ReScraper is a 0.6B refiner that replaces the heuristic HTML-to-text stack of a pretraining data pipeline — a rule-based scraper followed by rule-based cleaning filters — with a single small language model. - **Data:** [`cx-cmu/ReScraper-Data`](https://huggingface.co/datasets/cx-cmu/ReScraper-Data) - **Code:** [`cxcscmu/ReScraper`](https://github.com/cxcscmu/ReScraper) ## Contents - [What the model does](#what-the-model-does) - [Files](#files) - [Usage](#usage) - [Training](#training) - [License](#license) ## What the model does The model reads the visible text of a raw web page, rendered one block per line with line ids ``, and generates a short program: 1. An `` block of line removals (`rm 1-4`, `rm 6`, …) that strips headers, menus, boilerplate and other non-content lines. 2. One operation tag for the page: | Tag | Meaning | |---|---| | `` | keep the extracted text as it is | | `` | remove further noisy lines and spans inside the page | | `` | drop the page entirely | | `` | rewrite a poorly written but informative page | An executor applies the program to the page, so extraction and cleaning happen in one model instead of a cascade of separate stages. ## Files ``` models/ReScraper/ the refiner (Stage 2 checkpoint), in Hugging Face format ``` ## Usage ```python from vllm import LLM, SamplingParams llm = LLM(model="cx-cmu/ReScraper", dtype="bfloat16", max_model_len=32768) params = SamplingParams(temperature=1.0, top_p=1.0, max_tokens=3072) system = open("prompts/student_system_stage2.txt").read() # from the code repository page = " ...\n ...\n" # rendered page, one block per line out = llm.chat( [[{"role": "system", "content": system}, {"role": "user", "content": page}]], params, ) print(out[0].outputs[0].text) ``` The system prompt and the executor that applies the generated program are in the code repository, together with the full pipeline (rendering, inference, post-filtering, deduplication and tokenization). ## Training The model is trained in two supervised stages, distilling three teachers: a main-content extractor, a refinement teacher that decides the operation and edits the text, and a rewriting teacher for pages that are informative but poorly written. Both training sets are released in [`cx-cmu/ReScraper-Data`](https://huggingface.co/datasets/cx-cmu/ReScraper-Data). ## License Apache 2.0.