| <!DOCTYPE html> |
| <html lang="en"> |
| <head> |
| <meta charset="UTF-8"> |
| <meta name="viewport" content="width=device-width, initial-scale=1.0"> |
| <title>IfGPT DATASET Quality Components</title> |
| <style> |
| body { |
| font-family: -apple-system, BlinkMacSystemFont, "Segoe UI", Helvetica, Arial, sans-serif; |
| max-width: 900px; |
| margin: 2em auto; |
| padding: 0 1.5em; |
| line-height: 1.6; |
| color: #24292f; |
| background-color: #ffffff; |
| } |
| h1 { |
| border-bottom: 2px solid #d0d7de; |
| padding-bottom: 0.3em; |
| font-size: 2em; |
| } |
| h2 { |
| border-bottom: 1px solid #d0d7de; |
| padding-bottom: 0.3em; |
| margin-top: 1.8em; |
| font-size: 1.5em; |
| } |
| code { |
| background-color: #f6f8fa; |
| padding: 0.2em 0.4em; |
| border-radius: 4px; |
| font-family: "SF Mono", Menlo, Consolas, monospace; |
| font-size: 0.9em; |
| } |
| pre { |
| background-color: #f6f8fa; |
| padding: 1em; |
| border-radius: 6px; |
| overflow-x: auto; |
| font-family: "SF Mono", Menlo, Consolas, monospace; |
| font-size: 0.9em; |
| } |
| pre code { |
| background: none; |
| padding: 0; |
| } |
| ul { |
| padding-left: 1.5em; |
| } |
| li { |
| margin-bottom: 0.6em; |
| } |
| hr { |
| border: none; |
| border-top: 1px solid #d0d7de; |
| margin: 2em 0; |
| } |
| a { |
| color: #0969da; |
| text-decoration: none; |
| } |
| a:hover { |
| text-decoration: underline; |
| } |
| .footer { |
| font-size: 0.9em; |
| color: #57606a; |
| } |
| </style> |
| </head> |
| <body> |
|
|
| <h1>IfGPT DATASET Quality Components</h1> |
|
|
| <h2>Objectives of the project IfGPT</h2> |
| <p> |
| The <strong>IfGPT Quality Pipeline</strong> is developed within the project |
| <strong>IfGPT: Infrastructure for Fine-tuning Pre-trained Large Language Models</strong>, |
| which aims to establish a freely accessible infrastructure for the selection and |
| pre-processing of large datasets for Bulgarian as well as tailored data for particular |
| industries and fine-tuning suitable freely available large language models for specific purposes. |
| </p> |
|
|
| <h2>IfGPT Dataset Quality Pipeline</h2> |
| <p> |
| Modular Java pipeline to process and add new text documents to the IfGPT Dataset, |
| which includes cleaning, deduplication and quality evaluation of Bulgarian texts. |
| </p> |
| <p>The pipeline includes:</p> |
| <ul> |
| <li> |
| <strong>Source-specific extraction</strong> — metadata and plain text extracted from |
| heterogeneous corpora (MARCELL, CURLICAT, BulNC, Wikipedia), each handled by a |
| dedicated class that implements the shared <code>SourceProcessor</code> interface and |
| extends <code>BaseSourceProcessor</code>. |
| </li> |
| <li> |
| <strong>Sentence splitting</strong> — <code>BulgarianSentenceSplitter</code> wraps the |
| Apache OpenNLP Bulgarian UD sentence-detection model, splitting every document into a |
| sentence-per-line sidecar file used by all downstream stages. |
| </li> |
| <li> |
| <strong>Boilerplate cleaning</strong> — <code>FileCleanProcessor</code> learns |
| site-specific boilerplate from a sample directory (lines appearing in ≥ 50 % of files) |
| and removes them alongside hardcoded patterns for HTML tags, navigation menus, URLs, |
| cookie banners, and more. |
| </li> |
| <li> |
| <strong>MinHash / LSH deduplication</strong> — <code>DeduplicationProcessor</code> |
| builds a MinHash signature index over the full existing corpus and detects |
| near-duplicate sentences in the new batch (Jaccard ≥ 0.90), writing a ranked TSV |
| report and optionally removing duplicates. |
| </li> |
| <li> |
| <strong>Per-sentence PII scoring</strong> — <code>PIIDetector</code> runs every |
| sentence through the Phileas engine (names, emails, phone numbers, IBANs, IP |
| addresses, etc.) and stores the proportion of flagged tokens as the |
| <code>PersonallyIdentifiableInformation</code> vector in metadata. |
| </li> |
| <li> |
| <strong>Per-sentence bias scoring</strong> — <code>BiasAnalyser</code> matches tokens |
| against <code>BiasLexicon</code> (3 787-entry Bulgarian Bias Dictionary v4) to detect |
| signal–evaluator pairs across five categories (gender, race/ethnicity, religion, |
| disability, appearance), storing the per-sentence coverage ratio as the |
| <code>BiasedInformation</code> vector in metadata. |
| </li> |
| </ul> |
|
|
| <p> |
| The full schema is enforced by <code>DocumentMetadata</code> (15 mandatory + 8 optional |
| fields) and the complete flow is managed by <code>IfGPTPipeline</code>, with |
| <code>IfGPTDatasetProcessor</code> as the main entry point. |
| </p> |
|
|
| <pre><code>source processors → sentence split → clean → deduplication → PII → bias → counts → final structuring</code></pre> |
|
|
| <h2>License</h2> |
| <p>Creative Commons Attribution 4.0 International (CC-BY-4.0)</p> |
|
|
| <hr> |
|
|
| <p class="footer"> |
| This work is part of the project <strong>Infrastructure for Fine-tuning Pre-trained |
| Large Language Models</strong>, Grant Agreement No. ПВУ – 55 from 12.12.2024 |
| /BG-RRP-2.017-0030-C01/.<br> |
| <a href="https://ifgpt.dcl.bas.bg/en/">https://ifgpt.dcl.bas.bg/en/</a> |
| </p> |
|
|
| </body> |
| </html> |