You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Please provide your details and agree to the LICENSE [simpler version] to request access.

Log in or Sign Up to review the conditions and access this model content.

Indic Translate banner

Indic Translate

Task Languages License Blog GitHub


What is Indic Translate?

A translation model for English and 22 Indian languages, built for document-length text.

Indic Translate translates between English and all 22 languages of the Eighth Schedule of the Indian Constitution, in both directions. It is a single-purpose translation model: it takes one instruction and one piece of source text, and returns the translation. Nothing else.

Unlike sentence-level MT systems, Indic Translate is designed to take a whole document in one request, up to 32,768 tokens, and return a translation that preserves the document's paragraph structure, headings, lists, and inline markup.

This release adds three new capabilities: translating directly to and from Romanized script (writing an Indic language with Latin letters, the way people commonly type without an Indic keyboard), transliteration, converting between native and Romanized script without changing the language, and code-mixed translation, to and from text that mixes English and an Indic language within a sentence.


Model Summary

Model name Indic Translate
Repository bodhan-ai/indic-translate
Task Text-to-text translation (English ↔ 22 Indian languages)
Architecture Gemma4ForConditionalGeneration (Gemma 4 E4B)
Base model google/gemma-4-E4B-it
Total parameters 7.94 B
Precision bfloat16
Context window 131,072 tokens architecturally; 32,768 validated
Training data ~28.7 B tokens combined, roughly a 90:10 split between core translation data (sentence + document) and the Romanized-translation/transliteration data added for this release

Capabilities

Indic Translate supports

1. Sentence translation

Translate individual sentences and short text snippets quickly and accurately across 44 language directions: English ↔ 22 Indian languages.

2. Document translation

The model can translate an entire document in a single request, supporting inputs of up to 32,768 tokens. Formatting is preserved in the translation. This was verified across five types of documents:

Document class What is preserved
Markdown Heading levels, lists, blockquotes, tables, inline code, link targets and anchors, bare URLs
LaTeX Math environments, macros, equation and label structure
Tables and figures Row/column structure, numeric cells, units, figure and caption references
Code and docstrings Fenced code blocks byte-for-byte; prose, comments, and docstrings translated around them
Mixed-language documents Passages already in the target language are left alone rather than round-tripped

3. Romanized translation

Translate directly between English and Romanized Indic script, an Indic language written with Latin letters (e.g. aap kaise ho for Hindi), in both directions, across the same 22 languages.

4. Transliteration

Convert text between an Indic language's native script and its Romanized form, in either direction, without translating the meaning or changing the language.

5. Code-mixed translation

Translate to and from code-mixed text, a sentence that mixes English and an Indic language, the way people commonly type when switching between languages or without an Indic keyboard (e.g. "Yaar tomorrow ka meeting cancel ho gaya hai"). The model understands code-mixed input directly and can also produce code-mixed output on request.


Supported languages

Use the string in the Prompt name column as the target language. For the three multi-script languages, it also determines the output script.

Code Language Script Prompt name
eng_Latn English Latin English
asm_Beng Assamese Bengali–Assamese Assamese
ben_Beng Bengali Bengali Bengali
brx_Deva Bodo Devanagari Bodo
doi_Deva Dogri Devanagari Dogri
gom_Deva Konkani Devanagari Konkani
guj_Gujr Gujarati Gujarati Gujarati
hin_Deva Hindi Devanagari Hindi
kan_Knda Kannada Kannada Kannada
kas_Arab Kashmiri Perso-Arabic Kashmiri (Perso-Arabic script)
mai_Deva Maithili Devanagari Maithili
mal_Mlym Malayalam Malayalam Malayalam
mar_Deva Marathi Devanagari Marathi
mni_Beng Manipuri (Meitei) Bengali Manipuri (Bengali script)
mni_Mtei Manipuri (Meitei) Meetei Mayek Manipuri (Meitei script)
npi_Deva Nepali Devanagari Nepali
ory_Orya Odia Odia Odia
pan_Guru Punjabi Gurmukhi Punjabi
san_Deva Sanskrit Devanagari Sanskrit
sat_Olck Santali Ol Chiki Santali
snd_Arab Sindhi Perso-Arabic Sindhi (Perso-Arabic script)
snd_Deva Sindhi Devanagari Sindhi (Devanagari script)
tam_Taml Tamil Tamil Tamil
tel_Telu Telugu Telugu Telugu
urd_Arab Urdu Perso-Arabic Urdu

That is 22 languages plus English, across 25 language–script combinations.

Script defaults. For the three languages carried in more than one script, pass the qualified prompt name from the table to select the script you want. Where a bare name is used, the conventional defaults are Perso-Arabic for Kashmiri, Meetei Mayek for Manipuri, and Devanagari for Sindhi.

Romanized script. Every language above also has a Romanized form, requested the same way as a qualified script name: append (Roman) to the language name, e.g. Hindi (Roman). Use it as the target of a translation request (Translate the following text into Hindi (Roman):) or of a transliteration request (Transliterate the following text into Hindi (Roman):, or the reverse, into Hindi, to go back to native script). See Romanized translation & transliteration quality.

Code-mixed script. Every language above also has a code-mixed form, requested the same way: append (code-mixed) to the language name, e.g. Hindi (code-mixed). Use it as the target of a translation request to get output that mixes English and the target language within a sentence; code-mixed input needs no marker of its own.


Document-level quality

We evaluated 500 documents across all 22 languages (498 distinct documents, two reused across a pair of languages each; ~10,950 document–language cells scored per judge after quality-control exclusions), with each document scored on six criteria by two independent judges. The model is compared against Sarvam Translate, which also translates documents.

Automatic (dBLEU / dchrF++, document-level):

Metric Sarvam Translate This model
dBLEU 47.44 58.97
dchrF++ 57.96 69.44

LLM-judge (two independent judges, Gemini 3 Flash and GPT-5.6):

Measure Judge Sarvam Translate This model
Mean score (0–100) Gemini 3 Flash 63.99 82.49
Mean score (0–100) GPT-5.6 57.89 69.74
Share of documents that passed Gemini 3 Flash 29.9% 62.7%
Share of documents that passed GPT-5.6 16.1% 38.2%

Grouped bar charts comparing Sarvam Translate and this model on document-level LLM-judge
scores. Left panel: mean score out of 100 for each of two judges, Gemini 3 Flash and
GPT-5.6, this model scores 82.5 and 69.7 versus Sarvam's 64.0 and 57.9. Right panel: share
of documents that passed (scored at least 75 and kept both meaning and formatting), this
model passes 62.7% and 38.2% versus Sarvam's 29.9% and 16.1%. This model's bars are taller
than Sarvam's on every measure in both panels.

A document "passed" only if it scored at least 75 out of 100 and kept the meaning and kept the formatting, so it is a strict bar. Both judges agree this model clearly beats Sarvam Translate on every measure: mean score and pass rate, on both judges independently.

By criterion: Gemini 3 Flash judge, 1–5 scale
Criterion Sarvam Translate This model
Faithfulness 3.64 4.79
Fluency 4.11 4.85
Terminology 4.09 4.85
Format preservation 3.90 4.95
Protected-span handling 3.12 4.19
Style-specific quality 3.25 4.38

Translation quality

Evaluated on IN22-Gen. The model is evaluated on IN22-Gen, covering all 22 Indian languages in both directions, with 1,024 test sentences per translation direction. We use greedy decoding and report results using the chrF++ metric.

Direction chrF++
English → Indic 47.7
Indic → English 60.4

Average across all 22 languages.

By language

Language English → Indic Indic → English
Assamese 45.5 64.5
Bengali 48.9 62.1
Bodo 48.5 54.2
Dogri 54.5 66.9
Gujarati 52.3 64.9
Hindi 55.5 64.1
Kannada 51.1 63.2
Kashmiri 34.4 56.7
Konkani 43.8 54.6
Maithili 50.2 63.0
Malayalam 47.1 61.8
Manipuri 44.3 52.7
Marathi 50.2 62.3
Nepali 49.3 66.1
Odia 47.3 64.8
Punjabi 48.4 59.2
Sanskrit 36.0 53.6
Santali 39.2 47.3
Sindhi 37.5 55.5
Tamil 50.2 58.4
Telugu 51.3 62.7
Urdu 64.0 69.6

LLM-judge scores

chrF++ measures similarity to one particular reference translation. To measure quality directly, every English → Indic translation was additionally scored by Gemini 3.5 Flash acting as a judge, on two axes, each scored from 1 to 5:

  • Adequacy: is the meaning of the source preserved?
  • Fluency: does the output read naturally to a native speaker?

We perform reference-free evaluation. 22,528 sentences were scored (1,024 per language), with 15 scoring errors (n=22,513 scored).

Score (1–5)
Adequacy 4.90
Fluency 4.75

Average across all 22 languages.

Grouped vertical bar chart of adequacy and fluency for English to Indic translation across 22
languages, sorted by fluency from highest on the left to lowest on the right. Each language has a
blue adequacy column beside an orange fluency column, both starting at 1 and labelled with their
exact value. Adequacy ranges from 4.67 (Santali) to 4.97 (Maithili); fluency ranges
from 4.28 (Santali) to 4.92 (Maithili). The orange fluency column is shorter than the blue
adequacy column for every language, and the difference grows toward the right-hand side, where
Manipuri, Santali, Sanskrit, and Sindhi sit.

Fluency sits below adequacy in every one of the 22 languages. The two columns are close together on the left and separate visibly toward the right, where the model preserves the meaning but phrases it less naturally.

By language: adequacy and fluency
Language Adequacy Fluency
Assamese 4.91 4.82
Bengali 4.93 4.84
Bodo 4.91 4.81
Dogri 4.96 4.90
Gujarati 4.95 4.87
Hindi 4.95 4.82
Kannada 4.96 4.88
Kashmiri 4.89 4.75
Konkani 4.93 4.82
Maithili 4.97 4.92
Malayalam 4.93 4.86
Manipuri 4.69 4.43
Marathi 4.93 4.80
Nepali 4.92 4.80
Odia 4.92 4.83
Punjabi 4.91 4.80
Sanskrit 4.84 4.29
Santali 4.67 4.28
Sindhi 4.86 4.56
Tamil 4.93 4.79
Telugu 4.96 4.85
Urdu 4.94 4.82

Fluency scores lower than adequacy in every language (4.90 against 4.75 on average). The model preserves meaning more reliably than it produces natural phrasing, and the gap is widest in the lowest-resource languages, lowest fluency: Manipuri (4.43), Santali (4.28), Sanskrit (4.29), Sindhi (4.56), Kashmiri (4.75). Only the English → Indic direction was judged; the reverse direction is covered by chrF++ above.


How it compares

All numbers below come from the same IN22-Gen test set, the same 1,024 sentences per direction, and the same scoring code, so they can be compared directly.

Sentence-level chrF++

Each table is ranked on its own direction. This model places 3rd for English → Indic and 3rd for Indic → English. Rows marked with * cover only 16 of the 22 IN22-Gen languages (translategemma has no code path for Bodo, Dogri, Konkani, Maithili, Manipuri, or Santali) and are not directly comparable to the 22-language aggregates around them.

English → Indic

# Model chrF++
1 Sarvam Translate 48.9
2 IndicTrans2 48.6
3 Indic Translate 47.7
4 gemini-3.5-flash 40.3
5 translategemma-27b-it* 39.6
6 gemma-4-31B-it 39.2
7 translategemma-12b-it* 37.1
8 gemma-4-26B-a4b-it 36.2
9 gemma-3-27b-it 31.9
10 translategemma-4b-it* 31.2
11 gemma-3-12b-it 28.8
12 gemma-3-4b-it 22.9

Indic → English

# Model chrF++
1 gemini-3.5-flash 66.6
2 IndicTrans2 62.8
3 Indic Translate 60.4
4 gemma-4-31B-it 59.6
5 translategemma-27b-it* 58.5
6 gemma-4-26B-a4b-it 56.6
7 gemma-3-27b-it 55.5
8 translategemma-12b-it* 55.0
9 gemma-3-12b-it 53.2
10 Sarvam Translate 52.4
11 translategemma-4b-it* 49.2
12 gemma-3-4b-it 47.1
By language: this model, Sarvam Translate and IndicTrans2
Language v4 En→Indic Sarvam IndicTrans2 v4 Indic→En Sarvam IndicTrans2
Assamese 45.5 46.4 47.1 64.5 54.1 65.8
Bengali 48.9 49.4 51.8 62.1 53.1 63.2
Bodo 48.5 48.9 47.8 54.2 48.9 62.1
Dogri 54.5 58.2 57.8 66.9 56.3 72.6
Gujarati 52.3 53.1 53.5 64.9 55.2 66.5
Hindi 55.5 56.4 56.7 64.1 56.1 65.4
Kannada 51.1 52.2 51.0 63.2 53.4 64.2
Kashmiri 34.4 35.1 40.2 56.7 50.1 60.4
Konkani 43.8 44.9 45.2 54.6 48.2 59.2
Maithili 50.2 50.5 48.7 63.0 55.2 64.8
Malayalam 47.1 47.7 50.9 61.8 52.5 64.5
Manipuri 44.3 45.5 44.6 52.7 47.3 57.9
Marathi 50.2 51.3 51.0 62.3 54.3 63.7
Nepali 49.3 49.9 49.0 66.1 56.9 67.7
Odia 47.3 48.2 43.9 64.8 54.6 66.2
Punjabi 48.4 48.3 50.6 59.2 53.6 63.4
Sanskrit 36.0 38.1 38.8 53.6 47.5 54.8
Santali 39.2 40.9 33.4 47.3 43.4 45.3
Sindhi 37.5 39.1 36.6 55.5 49.0 57.3
Tamil 50.2 50.9 49.5 58.4 50.1 59.8
Telugu 51.3 51.6 52.4 62.7 53.1 64.8
Urdu 64.0 65.3 68.2 69.6 58.6 73.0

Romanized translation & transliteration quality

Head-to-head against Sarvam Translate, on a held-out sentence-level test split (22,600+ sentences per direction across 22–23 languages), the only split Sarvam Translate has been scored on for this capability:

Romanized translation (chrF++):

Direction This model Sarvam Translate
English → Romanized Indic 41.1 2.2
Romanized Indic → English 64.9 32.9

Transliteration (character error rate, lower is better):

Direction This model (CER) Sarvam Translate (CER)
Native script → Romanized 9.7% 81.6%
Romanized → native script 8.7% 62.0%

Sarvam Translate can partially read Romanized Indic input but essentially cannot write Romanized output (2.2 chrF++, 81.6% character error rate), these are new capabilities for this model line, not an incremental gain over an existing one.

On the IN22-Gen test set (the same 1,024 sentences per language used in Translation quality above, no Sarvam Translate comparison available on this split): English → Romanized Indic 36.6 chrF++, Romanized Indic → English 52.6 chrF++, native → Romanized transliteration 8.0% CER, Romanized → native transliteration 7.8% CER. These are a harder, more conservative estimate than the head-to-head numbers above, same model, different (harder) sentences.

On whole 32k-token documents, scored across every document including generation failures: English → Romanized 49.0 chrF++, Romanized → English 81.6 chrF++, native → Romanized transliteration 77.2 chrF++ (14.9% CER), Romanized → native transliteration 78.1 chrF++ (20.9% CER). Between 1.7% and 6.0% of documents per direction hit a generation failure (truncated or malformed output) at this length; scoring only the documents that generated cleanly raises these to 58.9 / 82.6 / 81.5 (9.1% CER) / 87.0 (9.8% CER) chrF++. Document-length Romanized translation and transliteration work reliably, but, unlike the sentence-level numbers above, a meaningful minority of full-document generations in this new capability need a retry or a shorter chunk size.


Human evaluation

Coming soon. Both tables below evaluate the previous release, not the checkpoint described elsewhere on this page, human evaluation has not yet been run for the current release. They're kept here, clearly labeled, as the most recent reference point available while a fresh pass is pending. Every other section on this card (automatic and LLM-judge scoring, sentence and document level) reflects the current release.

Sentence-level

The model was compared with Sarvam Translate on 100 sentences per language across all 22 languages (2,200 sentences in total). For each pair, evaluators selected a winner or marked both translations as good or poor.

Verdict Sentences Share
Previous release preferred 541 24.8%
Sarvam Translate preferred 595 27.2%
Both good 758 34.7%
Both poor 290 13.3%

The two were rated equal on 48% of sentences. Counting only the sentences where a winner was picked, Sarvam was chosen 52.4% of the time and the previous release 47.6%, close, and behind. The previous release leads in 10 of the 22 languages, so which model wins depends on the language.

By language: human verdicts, 100 sentences each (previous release)
Language Previous release Sarvam Translate Both good Both poor
Urdu 60 26 13 1
Marathi 44 22 11 22
Manipuri 26 10 21 43
Punjabi 13 8 77 2
Hindi 19 15 63 3
Malayalam 21 17 35 27
Santali 19 15 64 2
Tamil 31 27 24 18
Dogri 38 36 14 12
Gujarati 25 23 46 6
Bengali 39 40 16 3
Kannada 21 22 38 17
Telugu 19 21 28 31
Odia 26 29 39 6
Nepali 24 29 36 11
Kashmiri 25 33 30 12
Maithili 12 27 44 17
Konkani 29 47 19 2
Assamese 17 36 38 7
Sanskrit 4 24 65 7
Bodo 21 48 13 17
Sindhi 8 40 24 24

Document-level

A separate document-level human evaluation, also comparing the previous release against Sarvam Translate, is under way: 1,275 of 2,200 tasks scored (58.0%) as of 2026-08-27, this section's most recent report. Treat these as provisional until a completed pass is available; no document-level human evaluation has been run yet for the current release either.

Dimension (0–5, except Overall) Sarvam Translate Previous release Δ
Faithfulness 3.67 3.92 +0.25
Fluency 3.61 3.85 +0.25
Terminology 3.55 3.80 +0.25
Format preservation 4.11 4.37 +0.26
Protected spans 3.91 4.25 +0.34
Style-specific quality 3.71 3.96 +0.25
Overall (0–100) 75.3 80.6 +5.3

Example translations

Unedited model output, produced with greedy decoding using the prompt format described in the Prompting guide below.

Sentence translation

English → Hindi

Source Indus Valley Civilisation is known for its technological knowledge in a variety of fields.

Output सिंधु घाटी सभ्यता विभिन्न क्षेत्रों में अपने तकनीकी ज्ञान के लिए जानी जाती है।

English → Kannada

Source Hitopadesa is a book of worldly wisdom presented through the characters of birds, animals, and humans.

Output ಹಿಟೋಪದೇಶವು ಪಕ್ಷಿಗಳು, ಪ್ರಾಣಿಗಳು ಮತ್ತು ಮನುಷ್ಯರ ಪಾತ್ರಗಳ ಮೂಲಕ ಪ್ರಸ್ತುತಪಡಿಸಲಾದ ಲೌಕಿಕ ಜ್ಞಾನದ ಪುಸ್ತಕವಾಗಿದೆ.

English → Tamil

Source The Bagh Caves, consisting of Buddhist mural paintings, are located 97 km from the Dhar district of Madhya Pradesh.

Output பௌத்த சுவரோவியங்களைக் கொண்ட பாக் குகைகள் மத்தியப் பிரதேசத்தின் தார் மாவட்டத்திலிருந்து 97 கி.மீ தொலைவில் அமைந்துள்ளன.

English → Sanskrit

Source The purpose of the book appears to encourage proficiency in Sanskrit expression and the knowledge of wise behaviour.

Output पुस्तकस्य उद्देश्यं संस्कृतव्यञ्जने प्रवीणतां, बुद्धिमत्तापूर्णव्यवहारज्ञानं च प्रोत्साहयति इति प्रतीयते।

Hindi → English, the reverse direction, on formal register:

Source हाल के वर्षों में, रुपये की विनिमय दर तथा विदेशी मुद्रा आरक्षित भंडार की पर्याप्तता पर भिन्न-भिन्न दृष्टिकोण सामने आए हैं।

Output In recent years, there have been varying perspectives on the exchange rate of the rupee and the adequacy of foreign exchange reserves.

Note in the Tamil example that the numeral 97 and the unit are carried through rather than transliterated, and that named entities (Bagh, Dhar, Madhya Pradesh) are rendered in the target script rather than left in Latin.

Lower-resource languages: Kashmiri, Santali, Manipuri, Bodo, Maithili, Konkani

These are the languages with the least available training data, and the ones the Limitations section singles out for closer review. The output below is unedited, but it is illustrative rather than average: each example sits around the 90th percentile of its language's chrF++ distribution. For the figure that does describe typical quality, see the per-language table under Translation quality.

English → Kashmiri: Perso-Arabic script, right-to-left

Source Before 2014, 77 Passport Seva Kendras (PSK) were operational in the country.

Output 2014 برٛونٛہہ ٲسہِ ملکس منٛز 77 پاسپورٹ سیوا کیندر (پی ایس کے) کٲم کران۔

English → Santali: Ol Chiki script

Source The country has withstood the shocks from COVID-19 and the conflict in Ukraine.

Output ᱫᱤᱥᱚᱢ ᱫᱚ ᱠᱳᱵᱷᱤᱰᱼ᱑᱙ ᱟᱨ ᱤᱭᱩᱠᱨᱮᱱ ᱨᱮᱭᱟᱜ ᱞᱟᱹᱲᱦᱟᱹᱭ ᱨᱮᱭᱟᱜ ᱟᱱᱟᱴ ᱠᱚ ᱥᱟᱦᱟᱣ ᱟᱠᱟᱫᱟ ᱾

English → Manipuri: Meetei Mayek script

Source The first T-20 international match took place on August 5, 2004, between the women's teams of England and New Zealand.

Output ꯑꯍꯥꯟꯕ ꯇꯤ-꯲꯰ ꯏꯟꯇꯔꯅꯦꯁꯅꯦꯜ ꯃꯦꯆ ꯑꯁꯤ ꯏꯪ ꯲꯰꯰꯴ꯒꯤ ꯑꯣꯒꯁꯠ ꯵ꯗ ꯏꯪꯂꯦꯟꯗ ꯑꯃꯁꯨꯡ ꯅ꯭ꯌꯨ ꯖꯤꯂꯦꯟꯗꯒꯤ ꯅꯨꯄꯤꯁꯤꯡꯒꯤ ꯇꯤꯝꯁꯤꯡꯒꯤ ꯃꯔꯛꯇ ꯄꯥꯡꯊꯣꯛꯈꯤ꯫

English → Bodo: Devanagari script

Source And does 41-year-old Zlatan Ibrahimovic still have one big goal in him?

Output आरो 41 बोसोर बैसोनि ज्लाटान इब्राहिमेभिकनाव दासिमबो मोनसे गिदिर थांखि दं नामा?

English → Maithili: Devanagari script

Source According to a 2010 report, only about 2.1% (or 1,350,000 ha) of Afghanistan is forested.

Output २०१०क एकटा रिपोर्टक अनुसार अफगानिस्तानक मात्र २.१% (वा १,३५०,००० हेक्टेयर) भागमे जंगल अछि।

English → Konkani: Devanagari script

Source The game of Volleyball was invented by William G. Morgan in the year 1895 in Holyoke, Massachusetts by YMCA (USA).

Output व्हॉलीबॉल खेळाचो सोद 1895 वर्सा मॅसॅच्युसेट्सच्या होलोके हांगा वाय.एम.सी.ए. (यू.एस.ए.) न विल्यम जी. मॉर्गन हाणें लायिल्लो.

The reverse direction is consistently stronger:

Kashmiri → English

Source گٲمی عدالتہٕ، یتھ کٔنہِ زَن لوک عدالت (عوامٕی عدالت) یا نیایہ پنچایت (گامَن ہُنٛد انصاف)، چھِ تَضادٗك مُتبادل حل پیش کران چھِ ۔

Output Village courts, such as Lok Adalat (people's court) or Nyaya Panchayat (village justice), offer alternative dispute resolution.

Santali → English

Source ᱱᱮᱥᱱᱮᱞ ᱯᱷᱤᱞᱢ ᱟᱨᱠᱟᱭᱤᱵᱷᱥ ᱚᱯᱷ ᱤᱱᱰᱤᱭᱟ , ᱯᱩᱱᱮ ᱱᱚᱶᱟ ᱠᱤᱱ ᱵᱟᱨᱭᱟ ᱢᱤᱫ ᱠᱟᱛᱮ ᱱᱚᱶᱟ ᱠᱳᱨᱥ ᱫᱚᱠᱚ ᱥᱟᱯᱲᱟᱣᱞᱮᱫ ᱛᱟᱦᱮᱫᱽ ᱾

Output The National Film Archives of India, Pune, organized this course in collaboration with them.

Manipuri → English

Source ꯔꯤꯖꯅꯦꯜ ꯒ꯭ꯔꯤꯗꯁꯤꯡꯒꯤ ꯑꯍꯥꯟꯕ ꯑꯃꯒ-ꯑꯃꯒ ꯁꯝꯅꯕꯗꯨ ꯏꯪ ꯱꯹꯹꯱ꯒꯤ ꯑꯣꯛꯇꯣꯕꯔꯗ, ꯅꯣꯔꯊ ꯏꯁꯇꯔꯟ ꯑꯃꯗꯤ ꯏꯁꯇꯔꯟ ꯒ꯭ꯔꯤꯗꯁꯤꯡ ꯑꯃꯒ-ꯑꯃꯒ ꯁꯝꯅꯔꯕꯗ, ꯂꯤꯡꯈꯠꯈꯤ꯫

Output The first of the regional grids was established in October 1991, when the North Eastern and Eastern grids were merged.

Numerals are rendered in the target script's own digits for some languages and in Latin digits for others. Bengali, Santali, Manipuri, Maithili and Sanskrit output native digits (᱒᱐᱑᱓, ꯲꯰꯰꯴, २०१०); Hindi, Tamil, Telugu, Urdu, Bodo, Dogri, Konkani, Kashmiri and Sindhi keep Latin digits. Kannada is inconsistent and produces both. If a downstream system parses numbers out of the translation, normalise digits before doing so.

Document translation

The document below is a 4,196-character policy note containing a heading hierarchy, numbered and bulleted lists, a blockquote, HTML anchors, inline code, a Markdown link and a bare URL. Each translation was produced from a single request containing the entire document, and the output is unedited.

Source: English
# Data Retention and Deletion Policy (Draft)

## Summary

This policy note outlines the organization's approach to data retention, periodic review, and secure deletion. It is intended to ensure compliance with applicable laws, protect individual privacy, and minimize unnecessary storage of historical records. The policy applies to electronic and physical records created, received, or maintained by the organization in the course of business.

<a name="scope"></a>
## Scope

The scope of this policy includes customer records, employee records, transactional logs, backups, and analytics data. It excludes paper records governed by specialized archive agreements unless specifically referenced. For an itemized schedule, see the [Retention Schedule](https://company.example/policies/retention).

> Principle: retain only what is necessary, for as long as required, and dispose of it securely when it is no longer needed.

## Definitions

- `PII`: personally identifiable information, including names, email addresses, national identifiers, and equivalent identifiers.
- Retention period: the required time for which a given record must be kept before deletion.
- Archival storage: a lower-cost storage tier where records are preserved beyond active business use but prior to final deletion.

## Retention Schedule (high level)

1. Customer account data: retain for the life of the account plus 2 years for dispute resolution.
2. Financial and tax records: retain for a minimum of 7 years in accordance with accounting rules.
3. Employee personnel records: retain for 6 years after termination unless otherwise required by law.
4. System logs and telemetry: retain raw logs for 90 days, aggregated telemetry for 3 years.
5. Backups: retention follows the backup policy; deletions are governed by Section 3.2 and by disaster recovery requirements.

### Special categories

- Health data subject to HIPAA: retain and delete according to legal counsel guidance and secure deletion protocols.
- Data subject to litigation hold: suspend regular deletion until the hold is lifted.

## Data Deletion Procedures

<a name="deletion-procedure"></a>
When a record reaches the end of its retention period, it must be deleted or anonymized using the following steps:

1. Verify the retention metadata and confirm no holds or exceptions apply.
2. If eligible, move the record to the secure deletion queue.
3. Execute deletion using approved tools; example command for archival expiry might be `archive --expire 7y`.
4. Log the deletion event, including responsible operator ID, timestamp, and scope.

Operational notes:

- Deletion must be irreversible for `PII` and other sensitive fields.
- For distributed systems, ensure all replicas and tertiary storage tiers are included in the deletion workflow.

## Exceptions and Overrides

- Legal, regulatory, or investigatory holds override automated deletion. See Section 3.2 for hold procedures.
- Data required for open customer service tickets may be retained until case closure, then re-evaluated.

## Roles and Responsibilities

- Data Owners: assign retention classes to datasets and approve exceptions.
- IT Operations: run deletion workflows, maintain secure tooling, and verify completeness.
- Compliance Team: audit retention practices and maintain the authoritative Retention Schedule.

## Compliance, Audit, and Reporting

Periodic audits will validate adherence to retention classes, deletion logs, and hold procedures. Noncompliance will trigger remediation plans and may result in disciplinary action.

## FAQs

Q: How are retention periods determined?

A: Retention periods combine legal requirements, business need, and privacy considerations. When in doubt, consult the Compliance Team.

Q: Can archived data be restored after deletion?

A: No. Deletion is intended to be irreversible. For critical restorations, rely on backups that are within their retention window.

## Revision History

- Draft 1.0: initial policy note prepared for internal review.
- Next revision will include detailed retention matrix and technical appendices.

For additional guidance, contact the Data Governance Office or refer to the full schedule at the link above.
English → Hindi
# डेटा प्रतिधारण और विलोपन नीति (मसौदा)

## सारांश

यह नीति नोट डेटा प्रतिधारण, आवधिक समीक्षा और सुरक्षित विलोपन के लिए संगठन के दृष्टिकोण की रूपरेखा प्रस्तुत करता है। इसका उद्देश्य लागू कानूनों का अनुपालन सुनिश्चित करना, व्यक्तिगत गोपनीयता की रक्षा करना और ऐतिहासिक रिकॉर्ड के अनावश्यक भंडारण को कम करना है। यह नीति व्यवसाय के दौरान संगठन द्वारा बनाए गए, प्राप्त किए गए या बनाए रखे गए इलेक्ट्रॉनिक और भौतिक रिकॉर्ड पर लागू होती है।

<a name="scope"></a>
## दायरा

इस नीति के दायरे में ग्राहक रिकॉर्ड, कर्मचारी रिकॉर्ड, लेन-देन लॉग, बैकअप और एनालिटिक्स डेटा शामिल हैं। इसमें विशेष रूप से संदर्भित किए जाने तक विशेष अभिलेखीय समझौतों द्वारा शासित कागजी रिकॉर्ड शामिल नहीं हैं। एक विस्तृत अनुसूची के लिए, [प्रतिधारण अनुसूची](https://company.example/policies/retention) देखें।

> सिद्धांत: केवल वही बनाए रखें जो आवश्यक है, जितनी देर तक आवश्यक हो, और जब इसकी आवश्यकता न हो तो इसे सुरक्षित रूप से निपटा दें।

## परिभाषाएँ

- `PII`: व्यक्तिगत रूप से पहचान योग्य जानकारी, जिसमें नाम, ईमेल पते, राष्ट्रीय पहचानकर्ता और समकक्ष पहचानकर्ता शामिल हैं।
- प्रतिधारण अवधि: वह आवश्यक समय जिसके लिए विलोपन से पहले किसी दिए गए रिकॉर्ड को रखा जाना चाहिए।
- अभिलेखीय भंडारण: एक कम लागत वाला भंडारण स्तर जहाँ रिकॉर्ड सक्रिय व्यावसायिक उपयोग से परे लेकिन अंतिम विलोपन से पहले संरक्षित किए जाते हैं।

## प्रतिधारण अनुसूची (उच्च स्तर)

1. ग्राहक खाता डेटा: खाते के जीवनकाल के लिए और विवाद समाधान के लिए 2 वर्ष के लिए बनाए रखें।
2. वित्तीय और कर रिकॉर्ड: लेखांकन नियमों के अनुसार न्यूनतम 7 वर्षों के लिए बनाए रखें।
3. कर्मचारी कार्मिक रिकॉर्ड: कानून द्वारा अन्यथा आवश्यक होने तक, समाप्ति के बाद 6 वर्षों के लिए बनाए रखें।
4. सिस्टम लॉग और टेलीमेट्री: कच्चे लॉग 90 दिनों के लिए, एकत्रित टेलीमेट्री 3 वर्षों के लिए बनाए रखें।
5. बैकअप: प्रतिधारण बैकअप नीति का पालन करता है; विलोपन अनुभाग 3.2 और आपदा रिकवरी आवश्यकताओं द्वारा शासित होते हैं।

### विशेष श्रेणियां

- HIPAA के अधीन स्वास्थ्य डेटा: कानूनी परामर्श मार्गदर्शन और सुरक्षित विलोपन प्रोटोकॉल के अनुसार बनाए रखें और हटाएं।
- मुकदमेबाजी होल्ड के अधीन डेटा: होल्ड हटाए जाने तक नियमित विलोपन को निलंबित करें।

## डेटा विलोपन प्रक्रियाएं

<a name="deletion-procedure"></a>
जब कोई रिकॉर्ड अपनी प्रतिधारण अवधि के अंत तक पहुंच जाता है, तो उसे निम्नलिखित चरणों का उपयोग करके हटा दिया जाना चाहिए या गुमनाम कर दिया जाना चाहिए:

1. प्रतिधारण मेटाडेटा को सत्यापित करें और पुष्टि करें कि कोई होल्ड या अपवाद लागू नहीं होता है।
2. यदि योग्य है, तो रिकॉर्ड को सुरक्षित विलोपन कतार में ले जाएं।
3. अनुमोदित उपकरणों का उपयोग करके विलोपन निष्पादित करें; अभिलेखीय समाप्ति के लिए उदाहरण कमांड `archive --expire 7y` हो सकता है।
4. जिम्मेदार ऑपरेटर आईडी, टाइमस्टैम्प और दायरे सहित विलोपन घटना को लॉग करें।

परिचालन नोट्स:

- `PII` और अन्य संवेदनशील फ़ील्ड के लिए विलोपन अपरिवर्तनीय होना चाहिए।
- वितरित प्रणालियों के लिए, सुनिश्चित करें कि सभी प्रतिकृतियां और तृतीयक भंडारण स्तर विलोपन वर्कफ़्लो में शामिल हैं।

## अपवाद और ओवरराइड

- कानूनी, नियामक या जांच होल्ड स्वचालित विलोपन को ओवरराइड करते हैं। होल्ड प्रक्रियाओं के लिए अनुभाग 3.2 देखें।
- खुली ग्राहक सेवा टिकटों के लिए आवश्यक डेटा को मामले के बंद होने तक बनाए रखा जा सकता है, फिर पुनर्मूल्यांकन किया जा सकता है।

## भूमिकाएं और जिम्मेदारियां

- डेटा मालिक: डेटासेट को प्रतिधारण वर्ग असाइन करें और अपवादों को मंजूरी दें।
- आईटी संचालन: विलोपन वर्कफ़्लो चलाएं, सुरक्षित टूलिंग बनाए रखें और पूर्णता सत्यापित करें।
- अनुपालन टीम: प्रतिधारण प्रथाओं का ऑडिट करें और आधिकारिक प्रतिधारण अनुसूची बनाए रखें।

## अनुपालन, ऑडिट और रिपोर्टिंग

आवधिक ऑडिट प्रतिधारण वर्गों, विलोपन लॉग और होल्ड प्रक्रियाओं के पालन को मान्य करेंगे। गैर-अनुपालन से उपचारात्मक योजनाएं शुरू होंगी और इसके परिणामस्वरूप अनुशासनात्मक कार्रवाई हो सकती है।

## अक्सर पूछे जाने वाले प्रश्न

प्रश्न: प्रतिधारण अवधि कैसे निर्धारित की जाती है?

उत्तर: प्रतिधारण अवधि कानूनी आवश्यकताओं, व्यावसायिक आवश्यकता और गोपनीयता विचारों को जोड़ती है। संदेह होने पर, अनुपालन टीम से परामर्श करें।

प्रश्न: क्या अभिलेखीय डेटा को विलोपन के बाद पुनर्स्थापित किया जा सकता है?

उत्तर: नहीं। विलोपन अपरिवर्तनीय होने का इरादा है। महत्वपूर्ण पुनर्स्थापनाओं के लिए, उन बैकअप पर भरोसा करें जो उनकी प्रतिधारण विंडो के भीतर हैं।

## संशोधन इतिहास

- मसौदा 1.0: आंतरिक समीक्षा के लिए तैयार किया गया प्रारंभिक नीति नोट।
- अगले संशोधन में विस्तृत प्रतिधारण मैट्रिक्स और तकनीकी परिशिष्ट शामिल होंगे।

अतिरिक्त मार्गदर्शन के लिए, डेटा गवर्नेंस कार्यालय से संपर्क करें या ऊपर दिए गए लिंक पर पूर्ण अनुसूची देखें।
English → Tamil
# தரவு தக்கவைப்பு மற்றும் நீக்குதல் கொள்கை (வரைவு)

## சுருக்கம்

இந்தக் கொள்கைக் குறிப்பு, தரவு தக்கவைப்பு, அவ்வப்போது மதிப்பாய்வு செய்தல் மற்றும் பாதுகாப்பான நீக்குதல் ஆகியவற்றில் நிறுவனத்தின் அணுகுமுறையை கோடிட்டுக் காட்டுகிறது. இது பொருந்தக்கூடிய சட்டங்களுக்கு இணங்குவதை உறுதி செய்வதற்கும், தனிநபர் தனியுரிமையைப் பாதுகாப்பதற்கும், வரலாற்றுப் பதிவுகளைத் தேவையற்ற முறையில் சேமிப்பதைத் குறைப்பதற்கும் நோக்கம் கொண்டுள்ளது. வணிக நடவடிக்கைகளின் போது நிறுவனத்தால் உருவாக்கப்பட்ட, பெறப்பட்ட அல்லது பராமரிக்கப்படும் மின்னணு மற்றும் பௌதீகப் பதிவுகளுக்கு இந்தக் கொள்கை பொருந்தும்.

<a name="scope"></a>
## வரம்பு

இந்தக் கொள்கையின் வரம்பில் வாடிக்கையாளர் பதிவுகள், பணியாளர் பதிவுகள், பரிவர்த்தனைப் பதிவுகள், காப்புப் பிரதிகள் மற்றும் பகுப்பாய்வுத் தரவுகள் ஆகியவை அடங்கும். குறிப்பாகக் குறிப்பிடப்படாவிட்டால், சிறப்பு ஆவணக் காப்பக ஒப்பந்தங்களால் நிர்வகிக்கப்படும் காகிதப் பதிவுகள் இதில் அடங்காது. விரிவான அட்டவணைக்கு, [தக்கவைப்பு அட்டவணை](https://company.example/policies/retention) என்பதைப் பார்க்கவும்.

> கொள்கை: தேவைப்படும் வரை, தேவையானவற்றை மட்டுமே தக்கவைத்துக் கொள்ளவும், மேலும் அவை இனி தேவையில்லை என்றதும் பாதுகாப்பாக அவற்றை அப்புறப்படுத்தவும்.

## வரையறைகள்

- `PII`: பெயர்கள், மின்னஞ்சல் முகவரிகள், தேசிய அடையாளங்காட்டிகள் மற்றும் அதற்கு இணையான அடையாளங்காட்டிகள் உள்ளிட்ட தனிநபர் அடையாளம் காணக்கூடிய தகவல்.
- தக்கவைப்பு காலம்: ஒரு குறிப்பிட்ட பதிவை நீக்குவதற்கு முன்பு எவ்வளவு காலம் வைத்திருக்க வேண்டும் என்பதற்கான தேவையான நேரம்.
- ஆவணக் காப்பக சேமிப்பு: பதிவுகள் செயலில் உள்ள வணிகப் பயன்பாட்டிற்குப் பிறகும், ஆனால் இறுதி நீக்கத்திற்கு முன்பும் பாதுகாக்கப்படும் குறைந்த செலவுடைய சேமிப்பக அடுக்கு.

## தக்கவைப்பு அட்டவணை (உயர் நிலை)

1. வாடிக்கையாளர் கணக்குத் தரவு: கணக்கின் ஆயுட்காலம் மற்றும் தகராறு தீர்வுக்காகக் கூடுதலாக 2 ஆண்டுகள் தக்கவைத்துக் கொள்ளவும்.
2. நிதி மற்றும் வரிப் பதிவுகள்: கணக்கியல் விதிகளின்படி குறைந்தபட்சம் 7 ஆண்டுகள் தக்கவைத்துக் கொள்ளவும்.
3. பணியாளர் ஆவணப் பதிவுகள்: சட்டத்தால் வேறுவிதமாகத் தேவைப்பட்டால் ஒழிய, பணிநீக்கம் செய்யப்பட்ட பிறகு 6 ஆண்டுகள் தக்கவைத்துக் கொள்ளவும்.
4. கணினிப் பதிவுகள் மற்றும் தொலை அளவீடுகள்: மூலப் பதிவுகளை 90 நாட்களுக்கும், ஒருங்கிணைக்கப்பட்ட தொலை அளவீடுகளை 3 ஆண்டுகளுக்கும் தக்கவைத்துக் கொள்ளவும்.
5. காப்புப் பிரதிகள்: காப்புப் பிரதிகளுக்கான கொள்கையைப் பின்பற்றி தக்கவைப்பு மேற்கொள்ளப்படும்; நீக்குதல்கள் பிரிவு 3.2 மற்றும் பேரிடர் மீட்புத் தேவைகளால் நிர்வகிக்கப்படும்.

### சிறப்புப் பிரிவுகள்

- HIPAA-க்கு உட்பட்ட சுகாதாரத் தரவு: சட்ட ஆலோசகரின் வழிகாட்டுதல் மற்றும் பாதுகாப்பான நீக்குதல் நெறிமுறைகளின்படி தக்கவைத்துக் கொள்ளவும் மற்றும் நீக்கவும்.
- வழக்குத் தொடரப்படுவதற்கான தரவு: வழக்குத் தடை நீக்கப்படும் வரை வழக்கமான நீக்கத்தை நிறுத்தி வைக்கவும்.

## தரவு நீக்குதல் நடைமுறைகள்

<a name="deletion-procedure"></a>
ஒரு பதிவு அதன் தக்கவைப்பு காலத்தின் முடிவை அடையும்போது, பின்வரும் படிகளைப் பயன்படுத்தி அது நீக்கப்பட வேண்டும் அல்லது அநாமதேயமாக்கப்பட வேண்டும்:

1. தக்கவைப்பு மெட்டாடேட்டாவைச் சரிபார்த்து, எந்தத் தடைகள் அல்லது விதிவிலக்குகளும் பொருந்தாது என்பதை உறுதிப்படுத்தவும்.
2. தகுதியுடையதாக இருந்தால், பதிவை பாதுகாப்பான நீக்குதல் வரிசைக்கு நகர்த்தவும்.
3. அங்கீகரிக்கப்பட்ட கருவிகளைப் பயன்படுத்தி நீக்குதலைச் செயல்படுத்தவும்; ஆவணக் காப்பக காலாவதிக்கு எடுத்துக்காட்டாக `archive --expire 7y` என்ற கட்டளை இருக்கலாம்.
4. பொறுப்பான ஆபரேட்டர் ID, நேர முத்திரை மற்றும் வரம்பு உள்ளிட்ட நீக்குதல் நிகழ்வைப் பதிவு செய்யவும்.

செயல்பாட்டு குறிப்புகள்:

- `PII` மற்றும் பிற முக்கியமான புலங்களுக்கு நீக்குதல் மீளமுடியாததாக இருக்க வேண்டும்.
- விநியோகிக்கப்பட்ட அமைப்புகளுக்கு, அனைத்து நகல்களும் மற்றும் மூன்றாம் நிலை சேமிப்பக அடுக்குகளும் நீக்குதல் பணிப்பாய்வில் சேர்க்கப்பட்டுள்ளதா என்பதை உறுதிப்படுத்தவும்.

## விதிவிலக்குகள் மற்றும் மேலதிகரிமைகள்

- சட்ட, ஒழுங்குமுறை அல்லது புலனாய்வுத் தடைகள் தானியங்கி நீக்குதலை மீறுகின்றன. தடை நடைமுறைகளுக்குப் பிரிவு 3.2-ஐப் பார்க்கவும்.
- திறந்த வாடிக்கையாளர் சேவை டிக்கெட்டுகளுக்குத் தேவையான தரவு, வழக்கு முடிவடையும் வரை தக்கவைக்கப்படலாம், பின்னர் மறு மதிப்பீடு செய்யப்படும்.

## பாத்திரங்கள் மற்றும் பொறுப்புகள்

- தரவு உரிமையாளர்கள்: தரவுத்தொகுப்புகளுக்குத் தக்கவைப்பு வகுப்புகளை ஒதுக்குதல் மற்றும் விதிவிலக்குகளை அங்கீகரித்தல்.
- IT செயல்பாடுகள்: நீக்குதல் பணிப்பாய்வுகளை இயக்குதல், பாதுகாப்பான கருவிகளைப் பராமரித்தல் மற்றும் முழுமையைச் சரிபார்த்தல்.
- இணக்கக் குழு: தக்கவைப்பு நடைமுறைகளைத் தணிக்கை செய்தல் மற்றும் அதிகாரப்பூர்வமான தக்கவைப்பு அட்டவணையைப் பராமரித்தல்.

## இணக்கம், தணிக்கை மற்றும் அறிக்கை

அவ்வப்போது செய்யப்படும் தணிக்கைகள் தக்கவைப்பு வகுப்புகள், நீக்குதல் பதிவுகள் மற்றும் தடை நடைமுறைகளுக்கு இணங்குவதை உறுதி செய்யும். இணங்காதது திருத்தத் திட்டங்களைத் தூண்டும் மற்றும் ஒழுங்கு நடவடிக்கை எடுக்க வழிவகுக்கும்.

## அடிக்கடி கேட்கப்படும் கேள்விகள்

கே: தக்கவைப்பு காலங்கள் எவ்வாறு தீர்மானிக்கப்படுகின்றன?

ப: தக்கவைப்பு காலங்கள் சட்டத் தேவைகள், வணிகத் தேவை மற்றும் தனியுரிமை பரிசீலனைகளை ஒருங்கிணைக்கின்றன. சந்தேகம் இருந்தால், இணக்கக் குழுவைத் தொடர்பு கொள்ளவும்.

கே: நீக்கப்பட்ட பிறகு ஆவணக் காப்பகத் தரவை மீட்டெடுக்க முடியுமா?

ப: முடியாது. நீக்குதல் என்பது மீளமுடியாததாக இருக்க வேண்டும் என்பதே நோக்கம். முக்கியமான மீட்டெடுப்புகளுக்கு, அவற்றின் தக்கவைப்பு வரம்பிற்குள் இருக்கும் காப்புப் பிரதிகளையே நம்பியிருக்கவும்.

## திருத்த வரலாறு

- வரைவு 1.0: உள் மதிப்பாய்விற்காகத் தயாரிக்கப்பட்ட ஆரம்பக் கொள்கைக் குறிப்பு.
- அடுத்த திருத்தத்தில் விரிவான தக்கவைப்பு மேட்ரிக்ஸ் மற்றும் தொழில்நுட்ப இணைப்புகள் ஆகியவை அடங்கும்.

கூடுதல் வழிகாட்டுதலுக்கு, தரவு ஆளுகைக் அலுவலகத்தைத் தொடர்பு கொள்ளவும் அல்லது மேலே உள்ள இணைப்பில் முழுமையான அட்டவணையைப் பார்க்கவும்.
English → Bengali
# ডেটা সংরক্ষণ এবং মুছে ফেলার নীতি (খসড়া)

## সারসংক্ষেপ

এই নীতিমালায় ডেটা সংরক্ষণ, পর্যায়ক্রমিক পর্যালোচনা এবং সুরক্ষিতভাবে মুছে ফেলার ক্ষেত্রে সংস্থার দৃষ্টিভঙ্গি তুলে ধরা হয়েছে। এর উদ্দেশ্য হল প্রযোজ্য আইন মেনে চলা, ব্যক্তিগত গোপনীয়তা রক্ষা করা এবং ঐতিহাসিক রেকর্ডের অপ্রয়োজনীয় সংরক্ষণ কমানো। এই নীতিটি ব্যবসার প্রয়োজনে সংস্থার তৈরি, প্রাপ্ত বা রক্ষণাবেক্ষণ করা ইলেকট্রনিক এবং ফিজিক্যাল রেকর্ডের ক্ষেত্রে প্রযোজ্য।

<a name="scope"></a>
## পরিধি

এই নীতির পরিধির মধ্যে গ্রাহকের রেকর্ড, কর্মচারীর রেকর্ড, লেনদেনের লগ, ব্যাকআপ এবং অ্যানালিটিক্স ডেটা অন্তর্ভুক্ত রয়েছে। বিশেষ আর্কাইভ চুক্তির অধীনে থাকা কাগজের রেকর্ডগুলি বিশেষভাবে উল্লেখ না করা পর্যন্ত এর অন্তর্ভুক্ত নয়। একটি বিস্তারিত সময়সূচির জন্য, [সংরক্ষণ সময়সূচী](https://company.example/policies/retention) দেখুন।

> নীতি: শুধুমাত্র প্রয়োজনীয় ডেটা, যতক্ষণ প্রয়োজন, সংরক্ষণ করুন এবং যখন এর আর প্রয়োজন নেই তখন নিরাপদে এটি নিষ্পত্তি করুন।

## সংজ্ঞা

- `PII`: ব্যক্তিগতভাবে সনাক্তযোগ্য তথ্য, যার মধ্যে নাম, ইমেল ঠিকানা, জাতীয় শনাক্তকারী এবং সমতুল্য শনাক্তকারী অন্তর্ভুক্ত।
- সংরক্ষণ সময়কাল: মুছে ফেলার আগে একটি নির্দিষ্ট রেকর্ড কতক্ষণ ধরে রাখতে হবে তার প্রয়োজনীয় সময়।
- আর্কাইভাল স্টোরেজ: একটি কম খরচের স্টোরেজ স্তর যেখানে সক্রিয় ব্যবসায়িক ব্যবহারের পরেও চূড়ান্ত মুছে ফেলার আগে রেকর্ডগুলি সংরক্ষণ করা হয়।

## সংরক্ষণ সময়সূচী (উচ্চ স্তরের)

1. গ্রাহকের অ্যাকাউন্ট ডেটা: অ্যাকাউন্টের জীবনকাল এবং বিরোধ নিষ্পত্তির জন্য অতিরিক্ত ২ বছর সংরক্ষণ করুন।
2. আর্থিক এবং ট্যাক্স রেকর্ড: অ্যাকাউন্টিং নিয়ম অনুযায়ী কমপক্ষে ৭ বছর সংরক্ষণ করুন।
3. কর্মচারীর পার্সোনেল রেকর্ড: আইন দ্বারা অন্যথায় প্রয়োজন না হলে চাকরি থেকে বরখাস্ত হওয়ার পরে ৬ বছর সংরক্ষণ করুন।
4. সিস্টেম লগ এবং টেলিমেট্রি: ৯০ দিনের জন্য র লগ এবং ৩ বছরের জন্য একত্রিত টেলিমেট্রি সংরক্ষণ করুন।
5. ব্যাকআপ: ব্যাকআপ নীতি অনুসরণ করে ডেটা সংরক্ষণ করতে হবে; মুছে ফেলার বিষয়টি ৩.২ ধারা এবং দুর্যোগ পুনরুদ্ধার প্রয়োজনীয়তা দ্বারা নিয়ন্ত্রিত হবে।

### বিশেষ বিভাগ

- HIPAA-এর অধীন স্বাস্থ্য ডেটা: আইনি পরামর্শের নির্দেশনা এবং সুরক্ষিতভাবে মুছে ফেলার প্রোটোকল অনুযায়ী সংরক্ষণ এবং মুছে ফেলুন।
- লিটিগেশন হোল্ডের অধীন ডেটা: হোল্ড তুলে না নেওয়া পর্যন্ত নিয়মিত মুছে ফেলা স্থগিত করুন।

## ডেটা মুছে ফেলার পদ্ধতি

<a name="deletion-procedure"></a>
যখন কোনো রেকর্ড তার সংরক্ষণ সময়কালের শেষ পর্যায়ে পৌঁছায়, তখন নিম্নলিখিত পদক্ষেপগুলি ব্যবহার করে এটি মুছে ফেলতে বা বেনামী করতে হবে:

1. সংরক্ষণ মেটাডেটা যাচাই করুন এবং নিশ্চিত করুন যে কোনো হোল্ড বা ব্যতিক্রম প্রযোজ্য নয়।
2. যোগ্য হলে, রেকর্ডটিকে সুরক্ষিত মুছে ফেলার সারিতে সরিয়ে নিন।
3. অনুমোদিত সরঞ্জাম ব্যবহার করে মুছে ফেলার প্রক্রিয়াটি সম্পাদন করুন; আর্কাইভাল মেয়াদ শেষ হওয়ার জন্য উদাহরণস্বরূপ কমান্ড হতে পারে `archive --expire 7y`।
4. মুছে ফেলার ঘটনাটি লগ করুন, যার মধ্যে দায়িত্বপ্রাপ্ত অপারেটরের আইডি, টাইমস্ট্যাম্প এবং পরিধি অন্তর্ভুক্ত থাকবে।

কার্যকরী নোট:

- `PII` এবং অন্যান্য সংবেদনশীল ফিল্ডের জন্য মুছে ফেলা অবশ্যই অপরিবর্তনীয় হতে হবে।
- ডিস্ট্রিবিউটেড সিস্টেমের জন্য, নিশ্চিত করুন যে সমস্ত রেপ্লিকা এবং টারশিয়ারি স্টোরেজ স্তর মুছে ফেলার ওয়ার্কফ্লোতে অন্তর্ভুক্ত রয়েছে।

## ব্যতিক্রম এবং ওভাররাইড

- আইনি, নিয়ন্ত্রক বা তদন্তমূলক হোল্ড স্বয়ংক্রিয় মুছে ফেলার উপর ওভাররাইড করে। হোল্ড পদ্ধতিগুলির জন্য ৩.২ ধারা দেখুন।
- খোলা গ্রাহক পরিষেবা টিকিটের জন্য প্রয়োজনীয় ডেটা কেস বন্ধ না হওয়া পর্যন্ত সংরক্ষণ করা যেতে পারে, তারপর পুনরায় মূল্যায়ন করা হবে।

## ভূমিকা এবং দায়িত্ব

- ডেটা মালিক: ডেটাসেটে সংরক্ষণ ক্লাস নির্ধারণ করুন এবং ব্যতিক্রম অনুমোদন করুন।
- আইটি অপারেশনস: মুছে ফেলার ওয়ার্কফ্লো চালান, সুরক্ষিত সরঞ্জাম রক্ষণাবেক্ষণ করুন এবং সম্পূর্ণতা যাচাই করুন।
- কমপ্লায়েন্স টিম: সংরক্ষণ অনুশীলন নিরীক্ষণ করুন এবং প্রামাণিক সংরক্ষণ সময়সূচী বজায় রাখুন।

## কমপ্লায়েন্স, নিরীক্ষা এবং রিপোর্টিং

পর্যায়ক্রমিক নিরীক্ষা সংরক্ষণ ক্লাস, মুছে ফেলার লগ এবং হোল্ড পদ্ধতিগুলির আনুগত্য যাচাই করবে। অমান্য করলে প্রতিকার পরিকল্পনা শুরু হবে এবং এর ফলে শাস্তিমূলক ব্যবস্থা নেওয়া হতে পারে।

## প্রায়শ জিজ্ঞাসিত প্রশ্নাবলী

প্রশ্ন: সংরক্ষণ সময়কাল কীভাবে নির্ধারণ করা হয়?

উত্তর: সংরক্ষণ সময়কাল আইনি প্রয়োজনীয়তা, ব্যবসার প্রয়োজন এবং গোপনীয়তার বিবেচনার সমন্বয়ে গঠিত। সন্দেহ হলে, কমপ্লায়েন্স টিমের সাথে পরামর্শ করুন।

প্রশ্ন: মুছে ফেলার পরে আর্কাইভ করা ডেটা পুনরুদ্ধার করা যায় কি?

উত্তর: না। মুছে ফেলা অপরিবর্তনীয় করার উদ্দেশ্যে করা হয়েছে। গুরুত্বপূর্ণ পুনরুদ্ধারের জন্য, তাদের সংরক্ষণ উইন্ডোর মধ্যে থাকা ব্যাকআপের উপর নির্ভর করুন।

## সংশোধন ইতিহাস

- খসড়া ১.০: অভ্যন্তরীণ পর্যালোচনার জন্য প্রস্তুতকৃত প্রাথমিক নীতিমালায়।
- পরবর্তী সংশোধনে বিস্তারিত সংরক্ষণ ম্যাট্রিক্স এবং প্রযুক্তিগত পরিশিষ্ট অন্তর্ভুক্ত থাকবে।

অতিরিক্ত নির্দেশনার জন্য, ডেটা গভর্নেন্স অফিসের সাথে যোগাযোগ করুন অথবা উপরের লিঙ্কে সম্পূর্ণ সময়সূচী দেখুন।
English → Kashmiri: Perso-Arabic script, right-to-left
# ڈیٹا ریٹینشن تہٕ ڈیلیشن پالیسی (ڈرافٹ)

## خلاصہ

یہِ پالیسی نوٹ چھُ تنظیمہِ ہُنٛد ڈیٹا ریٹینشن، وقتاً فوقتاً جائزہٕ، تہٕ محفوظ ڈیلیشن خٲطرٕ نقطہٕ نظر بیان کران۔ اَمیُک مقصد چھُ لاگو قانونن ہُنٛد تعمیل یقیٖنی بناوُن، انفرادی رازداری ہُنٛد تحفظ کرُن، تہٕ تاریخی ریکارڈن ہُنٛد غیر ضروری سٹوریج کم کرُن۔ یہِ پالیسی چھِ تنظیمہِ دٔسۍ کرنہٕ آمٕتۍ، حٲصل کرنہٕ آمٕتۍ، یا برقرار تھاونہٕ آمٕتۍ الیکٹرانک تہٕ فزیکل ریکارڈن پؠٹھ لاگو گژھان۔

<a name="scope"></a>
## دائرہ کار

اَمہِ پالیسی ہٕنٛدِس دائرہ کارس منٛز چھِ کسٹمر ریکارڈ، ملازمن ہٕنٛدۍ ریکارڈ، ٹرانزیکشنل لاگ، بیک اپ، تہٕ اینالیٹکس ڈیٹا شٲمل۔ اَتھ منٛز چھِ نہٕ خصوصی آرکائیو معاہدن تحت یِنہٕ وٲلۍ کاغذی ریکارڈ شٲمل، ییٚتہِ تام نہٕ خاص طورس پؠٹھ حوالہٕ دِنہٕ آمُت آسہِ۔ تفصیلی شیڈول خٲطرٕ، وُچھِو [ریٹینشن شیڈول](https://company.example/policies/retention)۔

> اصول: صِرِف سُہو برقرار تھاوِو یُس ضروٗری چھُ، ییٚتہِ تام ضروٗرت آسہِ، تہٕ ییٚلہِ اَمیُک ضروٗرت نہٕ آسہِ تیٚلہِ کٔرِو اَتھ محفوظ طریٖقس پؠٹھ ضایع۔

## تعریفہٕ

- `PII`: ذاتی طورس پؠٹھ شناخت کرنہٕ یِنہٕ وٲلۍ معلومات، یَتھ منٛز ناو، ای میل ایڈریس، قومی شناخت کنندٕ، تہٕ مساوی شناخت کنندٕ شٲمل چھِ۔
- ریٹینشن پیریڈ: سُہ ضروٗری وقت یَتھ خٲطرٕ کُنہِ دِنہٕ آمٕتۍ ریکارڈس ڈیلیٹ کرنہٕ برٛونٛہہ تھاوُن چھُ ضروٗری۔
- آرکائیول سٹوریج: اَکھ کم لاگتھ وول سٹوریج ٹائر ییٚتہِ ریکارڈ فعال کاروباری استعمالہٕ کھۄتہٕ زیٛادٕ مگر حتمی ڈیلیشن برٛونٛہہ محفوظ تھاونہٕ چھِ یِوان۔

## ریٹینشن شیڈول (اعلیٰ سطح)

1. کسٹمر اکاؤنٹ ڈیٹا: اکاؤنٹ کِس پوٗرٕ وقتس تہٕ تنازعہٕ حل کرنہٕ خٲطرٕ 2 ؤریو تام برقرار تھاوِو۔
2. مالی تہٕ ٹیکس ریکارڈ: اکاؤنٹنگ اصولن ہٕنٛدِ مطابق کم از کم 7 ؤریو تام برقرار تھاوِو۔
3. ملازمن ہٕنٛدۍ پرسنل ریکارڈ: نوکری ختم کرنہٕ پتہٕ 6 ؤریو تام برقرار تھاوِو، ییٚتہِ تام نہٕ قانون کہِ تحت کُنہِ بیٚیہِ ضروٗرت آسہِ۔
4. سسٹم لاگ تہٕ ٹیلی میٹری: 90 دۄہن خٲطرٕ را لاگ برقرار تھاوِو، 3 ؤریو خٲطرٕ ایگریگیٹڈ ٹیلی میٹری۔
5. بیک اپ: ریٹینشن چھُ بیک اپ پالیسی ہٕنٛدِ مطابق آسان؛ ڈیلیشن چھِ سیکشن 3.2 تہٕ ڈیزاسٹر ریکوری ہٕنٛز ضروٗرتن تحت یِوان۔

### خاص زمرٕ

- HIPAA کِس تحت صحتُک ڈیٹا: قانونی مشورٕ تہٕ محفوظ ڈیلیشن پروٹوکولن ہٕنٛدِ ہدایتن مطٲبق برقرار تھاوِو تہٕ ڈیلیٹ کٔرِو۔
- لٹیگیشن ہولڈ کِس تحت ڈیٹا: ییٚتہِ تام ہولڈ ہٹاونہٕ نہٕ یِیہِ، تیٚتہِ تام کٔرِو باقٲعدٕ ڈیلیشن معطل۔

## ڈیٹا ڈیلیشن طریقہٕ کار

<a name="deletion-procedure"></a>
ییٚلہِ کُنہِ ریکارڈس پننِس ریٹینشن پیریڈُک اختتام چھُ گژھان، تیٚلہِ چھُ یہِ ڈیلیٹ کرُن یا مندرجہ ذیل قدمو سۭتۍ گمنام کرُن ضروٗری:

1. ریٹینشن میٹا ڈیٹإچ تصدیٖق کٔرِو تہٕ یقیٖنی بناوِو زِ کانٛہہ تہِ ہولڈ یا استثنیٰ چھُ نہٕ لاگو گژھان۔
2. اگر اہل آسہِ، تیٚلہِ کٔرِو ریکارڈ محفوظ ڈیلیشن کیو (queue) منٛز منتقل۔
3. منظور شُدٕ ٹولن ہُنٛد استعمال کٔرِتھ کٔرِو ڈیلیشن لاگو؛ آرکائیول ایکسپائری خٲطرٕ مثال کمانڈ ہیکہِ `archive --expire 7y` ٲسِتھ۔
4. ڈیلیشن ایونٹ لاگ کٔرِو، یَتھ منٛز ذمہ دار آپریٹر ID، ٹائم سٹیمپ، تہٕ دائرہ کار شٲمل چھُ۔

آپریشنل نوٹ:

- `PII` تہٕ باقٕے حساس فیلڈز خٲطرٕ گژھِ ڈیلیشن ناقابلِ واپسی آسٕنۍ۔
- ڈسٹریبیوٹڈ سسٹمن خٲطرٕ، یقیٖنی بناوِو زِ تمام ریپلیکا تہٕ ٹرشری سٹوریج ٹائر چھِ ڈیلیشن ورک فلوہس منٛز شٲمل۔

## استثنیٰ تہٕ اوور رائیڈز

- قانونی، ریگولیٹری، یا تحقیقاتی ہولڈز چھِ خودکار ڈیلیشنس اوور رائیڈ کران۔ ہولڈ طریقہٕ کار خٲطرٕ وُچھِو سیکشن 3.2۔
- کھۄلہِ کسٹمر سروس ٹکٹن خٲطرٕ ضروٗری ڈیٹا ہیکہِ کیس بند گژھنہٕ تام برقرار تھاونہٕ یِتھ، پتہٕ ییٚیہِ دوبارٕ جائزہٕ نِنہٕ۔

## کردار تہٕ ذمہ داریاں

- ڈیٹا اونرز: ڈیٹاسیٹن ریٹینشن کلاسز تفویض کران تہٕ استثنیٰ منظور کران۔
- IT آپریشنز: ڈیلیشن ورک فلو چلاوان، محفوظ ٹولنگ برقرار تھاوان، تہٕ مکملیتٕچ تصدیٖق کران۔
- کمپلائنس ٹیم: ریٹینشن پریکٹسز ہُنٛد آڈٹ کران تہٕ مستند ریٹینشن شیڈول برقرار تھاوان۔

## تعمیل، آڈٹ، تہٕ رپورٹنگ

وقتاً فوقتاً آڈٹ کرنہٕ سۭتۍ گژھِ ریٹینشن کلاسز، ڈیلیشن لاگ، تہٕ ہولڈ طریقہٕ کارن ہٕنٛز تعمیل ہٕنٛز تصدیٖق۔ عدم تعمیل سۭتۍ گژھن اصلاحی منصوبن ہٕنٛز شروٗعات تہٕ اَمیُک نتیجٕہ ہیکہِ تادیبی کاروٲیی ٲسِتھ۔

## FAQs

س: ریٹینشن پیریڈ کِتھ کٔنۍ چھِ طے یِوان کرنہٕ؟

ج: ریٹینشن پیریڈ چھِ قانونی ضروٗرتن، کاروباری ضروٗرتن، تہٕ رازداری ہٕنٛدین پہلووَن مِلاوان۔ ییٚلہِ شکھ آسہِ، تیٚلہِ کٔرِو کمپلائنس ٹیمس سۭتۍ مشورٕ۔

س: کیا ڈیلیشن پتہٕ ہیکہِ آرکائیو کرنہٕ آمُت ڈیٹا بحال کرنہٕ یِتھ؟

ج: نہ۔ ڈیلیشن چھُ ناقابلِ واپسی آسُن۔ اہم بحالی خٲطرٕ، تِمَن بیک اپن پؠٹھ بھروسہٕ کٔرِو یِم پننِس ریٹینشن ونڈوٗہس منٛز چھِ۔

## نظر ثانی ہٕنٛز تٲریٖخ

- ڈرافٹ 1.0: ابتدٲیی پالیسی نوٹ یُس اندرونی جٲیزٕ خٲطرٕ تیار کرنہٕ آو۔
- اَمہِ پتہٕ کِس نظر ثانی منٛز آسہِ تفصیلی ریٹینشن میٹرکس تہٕ تکنیکی اپینڈکس شٲمل۔

مزید رہنمٲیی خٲطرٕ، کٔرِو ڈیٹا گورننس آفسس سۭتۍ رابطہٕ یا وُچھِو ہیٚرِمہِ لنکس پؠٹھ پوٗرٕ شیڈول۔
English → Santali: Ol Chiki script
# ᱰᱮᱴᱟ ᱫᱚᱦᱚᱭ ᱟᱨ ᱢᱮᱴᱟᱣ ᱨᱮᱭᱟᱜ ᱱᱤᱛᱤ (ᱰᱨᱟᱯᱷᱴ)

## ᱥᱟᱨᱟᱝᱥ

ᱱᱚᱣᱟ ᱱᱤᱛᱤ ᱱᱚᱴ ᱫᱚ ᱰᱮᱴᱟ ᱫᱚᱦᱚᱭ, ᱚᱠᱛᱚ ᱚᱠᱛᱚ ᱨᱮ ᱧᱮᱞ ᱵᱤᱰᱟᱹᱣ, ᱟᱨ ᱥᱩᱨᱩᱠᱷᱤᱛ ᱢᱮᱴᱟᱣ ᱞᱟᱹᱜᱤᱫ ᱥᱚᱝᱜᱚᱴᱷᱚᱱ ᱨᱮᱭᱟᱜ ᱩᱯᱟᱹᱭ ᱮ ᱩᱫᱩᱜᱽ ᱮᱫᱟ ᱾ ᱱᱚᱣᱟ ᱫᱚ ᱞᱟᱹᱜᱩ ᱟᱠᱟᱱ ᱟᱹᱱ ᱠᱚ ᱥᱟᱶ ᱢᱤᱞᱟᱹᱣ, ᱦᱚᱲ ᱠᱚᱣᱟᱜ ᱜᱚᱯᱚᱱᱤᱭᱚᱛᱟ ᱨᱩᱠᱷᱤᱭᱟᱹ, ᱟᱨ ᱢᱟᱲᱟᱝ ᱨᱮᱭᱟᱜ ᱨᱮᱠᱚᱨᱰ ᱠᱚ ᱵᱟᱝ ᱞᱟᱹᱜᱛᱤᱭᱟᱱ ᱫᱚᱦᱚᱭ ᱠᱚᱢ ᱞᱟᱹᱜᱤᱫ ᱛᱮᱭᱟᱨ ᱟᱠᱟᱱᱟ ᱾ ᱱᱚᱣᱟ ᱱᱤᱛᱤ ᱫᱚ ᱵᱮᱯᱟᱨ ᱚᱠᱛᱚ ᱨᱮ ᱥᱚᱝᱜᱚᱴᱷᱚᱱ ᱦᱚᱛᱮᱛᱮ ᱛᱮᱭᱟᱨ ᱟᱠᱟᱱ, ᱧᱟᱢ ᱟᱠᱟᱱ, ᱥᱮ ᱫᱚᱦᱚ ᱟᱠᱟᱱ ᱤᱞᱮᱠᱴᱨᱚᱱᱤᱠ ᱟᱨ ᱯᱷᱤᱡᱤᱠᱟᱞ ᱨᱮᱠᱚᱨᱰ ᱠᱚ ᱞᱟᱹᱜᱤᱫ ᱞᱟᱹᱜᱩᱜᱼᱟ ᱾

<a name="scope"></a>
## ᱥᱠᱳᱯ

ᱱᱚᱣᱟ ᱱᱤᱛᱤ ᱨᱮᱭᱟᱜ ᱥᱠᱳᱯ ᱨᱮ ᱜᱨᱟᱦᱚᱠ ᱨᱮᱠᱚᱨᱰ, ᱠᱟᱹᱢᱤᱭᱟᱹ ᱨᱮᱠᱚᱨᱰ, ᱴᱨᱟᱱᱡᱮᱠᱥᱚᱱᱟᱞ ᱞᱚᱜᱽ, ᱵᱮᱠᱟᱯ, ᱟᱨ ᱮᱱᱟᱞᱟᱭᱴᱤᱠᱥ ᱰᱮᱴᱟ ᱥᱮᱞᱮᱫ ᱢᱮᱱᱟᱜᱼᱟ ᱾ ᱱᱚᱣᱟ ᱫᱚ ᱵᱤᱥᱮᱥ ᱟᱨᱠᱟᱭᱤᱵᱽ ᱪᱩᱠᱛᱤ ᱦᱚᱛᱮᱛᱮ ᱪᱟᱞᱟᱜ ᱠᱟᱜᱚᱡᱽ ᱨᱮᱠᱚᱨᱰ ᱠᱚ ᱵᱟᱭ ᱥᱮᱞᱮᱫ ᱮᱫᱟ ᱡᱩᱫᱤ ᱵᱤᱥᱮᱥ ᱞᱮᱠᱟᱛᱮ ᱵᱟᱝ ᱩᱫᱩᱜ ᱟᱠᱟᱱᱟ ᱾ ᱢᱤᱫᱴᱟᱝ ᱟᱭᱴᱮᱢᱟᱭᱡᱽᱰ ᱥᱮᱰᱭᱩᱞ ᱞᱟᱹᱜᱤᱫ, [ᱨᱤᱴᱮᱱᱥᱚᱱ ᱥᱮᱰᱭᱩᱞ](https://company.example/policies/retention) ᱧᱮᱞ ᱢᱮ ᱾

> ᱱᱤᱭᱚᱢ: ᱡᱟᱦᱟᱸ ᱞᱟᱹᱜᱛᱤᱭᱟᱱᱟ, ᱩᱱᱟᱹᱜ ᱫᱤᱱ ᱞᱟᱹᱜᱤᱫ ᱫᱚᱦᱚᱭ ᱢᱮ, ᱟᱨ ᱡᱚᱠᱷᱚᱱ ᱵᱟᱝ ᱞᱟᱹᱜᱛᱤᱭᱟᱱᱟ ᱩᱱ ᱡᱚᱠᱷᱚᱱ ᱥᱩᱨᱩᱠᱷᱤᱛ ᱞᱮᱠᱟᱛᱮ ᱢᱮᱴᱟᱣ ᱢᱮ ᱾

## ᱯᱚᱨᱤᱵᱷᱟᱥᱟ

- `PII`: ᱦᱚᱲ ᱠᱚᱣᱟᱜ ᱩᱯᱨᱩᱢ ᱧᱟᱢ ᱞᱟᱹᱜᱤᱫ ᱰᱮᱴᱟ, ᱡᱟᱦᱟᱸ ᱨᱮ ᱧᱩᱛᱩᱢ, ᱤᱢᱮᱞ ᱮᱰᱨᱮᱥ, ᱡᱟᱹᱛᱤᱭᱟᱹᱨᱤ ᱩᱯᱨᱩᱢ, ᱟᱨ ᱚᱱᱟ ᱞᱮᱠᱟᱱ ᱩᱯᱨᱩᱢ ᱠᱚ ᱥᱮᱞᱮᱫ ᱢᱮᱱᱟᱜᱼᱟ ᱾
- ᱫᱚᱦᱚᱭ ᱚᱠᱛᱚ: ᱢᱤᱫᱴᱟᱝ ᱨᱮᱠᱚᱨᱰ ᱢᱮᱴᱟᱣ ᱢᱟᱲᱟᱝ ᱨᱮ ᱛᱤᱱᱟᱹᱜ ᱚᱠᱛᱚ ᱫᱚᱦᱚᱭ ᱞᱟᱹᱜᱛᱤᱭᱟᱱᱟ ᱚᱱᱟ ᱚᱠᱛᱚ ᱾
- ᱟᱨᱠᱟᱭᱤᱵᱟᱞ ᱥᱴᱳᱨᱮᱡᱽ: ᱢᱤᱫᱴᱟᱝ ᱠᱚᱢ ᱠᱷᱚᱨᱪᱟ ᱨᱮᱭᱟᱜ ᱥᱴᱳᱨᱮᱡᱽ ᱴᱤᱭᱟᱨ ᱡᱟᱦᱟᱸ ᱨᱮ ᱨᱮᱠᱚᱨᱰ ᱠᱚ ᱠᱟᱹᱢᱤ ᱵᱮᱯᱟᱨ ᱵᱮᱵᱷᱟᱨ ᱛᱟᱭᱚᱢ ᱦᱚᱸ ᱢᱮᱱᱠᱷᱟᱱ ᱢᱩᱪᱟᱹᱫ ᱢᱮᱴᱟᱣ ᱢᱟᱲᱟᱝ ᱨᱮ ᱫᱚᱦᱚ ᱦᱩᱭᱩᱜᱼᱟ ᱾

## ᱨᱤᱴᱮᱱᱥᱚᱱ ᱥᱮᱰᱭᱩᱞ (ᱞᱟᱯᱷᱟᱝ ᱛᱷᱟᱹᱨ)

1. ᱜᱨᱟᱦᱚᱠ ᱮᱠᱟᱣᱩᱱᱴ ᱰᱮᱴᱟ: ᱮᱠᱟᱣᱩᱱᱴ ᱨᱮᱭᱟᱜ ᱡᱤᱣᱚᱱ ᱚᱠᱛᱚ ᱟᱨ ᱵᱤᱵᱟᱹᱫᱽ ᱥᱚᱞᱦᱮ ᱞᱟᱹᱜᱤᱫ 2 ᱥᱮᱨᱢᱟ ᱫᱚᱦᱚᱭ ᱢᱮ ᱾
2. ᱠᱟᱹᱣᱰᱤᱟᱹᱨᱤ ᱟᱨ ᱴᱮᱠᱥ ᱨᱮᱠᱚᱨᱰ: ᱮᱠᱟᱣᱩᱱᱴᱤᱝ ᱱᱤᱭᱚᱢ ᱞᱮᱠᱟᱛᱮ ᱠᱚᱢ ᱥᱮ ᱠᱚᱢ 7 ᱥᱮᱨᱢᱟ ᱫᱚᱦᱚᱭ ᱢᱮ ᱾
3. ᱠᱟᱹᱢᱤᱭᱟᱹ ᱯᱟᱨᱥᱚᱱᱮᱞ ᱨᱮᱠᱚᱨᱰ: ᱠᱟᱹᱢᱤ ᱪᱟᱵᱟ ᱛᱟᱭᱚᱢ 6 ᱥᱮᱨᱢᱟ ᱫᱚᱦᱚᱭ ᱢᱮ ᱡᱩᱫᱤ ᱟᱹᱱ ᱦᱚᱛᱮᱛᱮ ᱮᱴᱟᱜ ᱞᱟᱹᱜᱛᱤ ᱵᱟᱝ ᱦᱩᱭᱩᱜᱼᱟ ᱾
4. ᱥᱤᱥᱴᱚᱢ ᱞᱚᱜᱽ ᱟᱨ ᱴᱮᱞᱤᱢᱮᱴᱨᱤ: 90 ᱢᱟᱦᱟᱸ ᱞᱟᱹᱜᱤᱫ ᱨᱚ ᱞᱚᱜᱽ ᱫᱚᱦᱚᱭ ᱢᱮ, 3 ᱥᱮᱨᱢᱟ ᱞᱟᱹᱜᱤᱫ ᱮᱜᱽᱨᱤᱜᱮᱴᱮᱰ ᱴᱮᱞᱤᱢᱮᱴᱨᱤ ᱫᱚᱦᱚᱭ ᱢᱮ ᱾
5. ᱵᱮᱠᱟᱯ: ᱨᱤᱴᱮᱱᱥᱚᱱ ᱫᱚ ᱵᱮᱠᱟᱯ ᱱᱤᱛᱤ ᱞᱮᱠᱟᱛᱮ ᱪᱟᱞᱟᱜᱼᱟ; ᱢᱮᱴᱟᱣ ᱫᱚ ᱥᱮᱠᱥᱚᱱ 3.2 ᱟᱨ ᱰᱤᱡᱟᱥᱴᱟᱨ ᱨᱤᱠᱚᱵᱷᱟᱨᱤ ᱞᱟᱹᱜᱛᱤ ᱠᱚ ᱦᱚᱛᱮᱛᱮ ᱪᱟᱞᱟᱜᱼᱟ ᱾

### ᱵᱤᱥᱮᱥ ᱠᱮᱴᱮᱜᱚᱨᱤ ᱠᱚ

- HIPAA ᱨᱮᱭᱟᱜ ᱟᱹᱱ ᱞᱮᱠᱟᱛᱮ ᱦᱚᱲᱢᱚ ᱥᱟᱶᱟᱨ ᱰᱮᱴᱟ: ᱟᱹᱱ ᱥᱟᱞᱦᱟ ᱟᱨ ᱥᱩᱨᱩᱠᱷᱤᱛ ᱢᱮᱴᱟᱣ ᱯᱨᱚᱴᱚᱠᱚᱞ ᱞᱮᱠᱟᱛᱮ ᱫᱚᱦᱚᱭ ᱟᱨ ᱢᱮᱴᱟᱣ ᱢᱮ ᱾
- ᱞᱤᱴᱤᱜᱮᱥᱚᱱ ᱦᱚᱞᱰ ᱨᱮᱭᱟᱜ ᱟᱹᱱ ᱞᱮᱠᱟᱛᱮ ᱰᱮᱴᱟ: ᱦᱚᱞᱰ ᱵᱟᱝ ᱪᱟᱵᱟᱜ ᱫᱷᱟᱹᱵᱤᱡ ᱫᱤᱱᱟᱹᱢ ᱢᱮᱴᱟᱣ ᱫᱚ ᱵᱚᱸᱫᱽ ᱢᱮ ᱾

## ᱰᱮᱴᱟ ᱢᱮᱴᱟᱣ ᱨᱮᱭᱟᱜ ᱯᱨᱚᱠᱨᱤᱭᱟ ᱠᱚ

<a name="deletion-procedure"></a>
ᱡᱚᱠᱷᱚᱱ ᱢᱤᱫᱴᱟᱝ ᱨᱮᱠᱚᱨᱰ ᱟᱡᱟᱜ ᱫᱚᱦᱚᱭ ᱚᱠᱛᱚ ᱨᱮᱭᱟᱜ ᱢᱩᱪᱟᱹᱫ ᱨᱮ ᱥᱮᱴᱮᱨᱚᱜᱼᱟ, ᱚᱱᱟ ᱫᱚ ᱞᱟᱛᱟᱨ ᱨᱮᱭᱟᱜ ᱯᱟᱹᱣᱲᱤ ᱠᱚ ᱵᱮᱵᱷᱟᱨ ᱠᱟᱛᱮ ᱢᱮᱴᱟᱣ ᱥᱮ ᱮᱱᱚᱱᱤᱢᱟᱭᱤᱡᱽ ᱦᱩᱭᱩᱜᱼᱟ:

1. ᱨᱤᱴᱮᱱᱥᱚᱱ ᱢᱮᱴᱟᱰᱮᱴᱟ ᱧᱮᱞ ᱢᱮ ᱟᱨ ᱯᱩᱥᱴᱟᱹᱣ ᱢᱮ ᱡᱮ ᱡᱟᱦᱟᱸᱱ ᱦᱚᱞᱰ ᱥᱮ ᱵᱟᱫᱽ ᱵᱟᱝ ᱞᱟᱹᱜᱩᱜ ᱠᱟᱱᱟ ᱾
2. ᱡᱩᱫᱤ ᱞᱟᱹᱜᱛᱤᱭᱟᱱᱟ, ᱨᱮᱠᱚᱨᱰ ᱫᱚ ᱥᱩᱨᱩᱠᱷᱤᱛ ᱢᱮᱴᱟᱣ ᱠᱤᱭᱩ ᱨᱮ ᱤᱫᱤ ᱢᱮ ᱾
3. ᱢᱟᱱᱟᱣ ᱟᱠᱟᱱ ᱴᱩᱞ ᱠᱚ ᱵᱮᱵᱷᱟᱨ ᱠᱟᱛᱮ ᱢᱮᱴᱟᱣ ᱠᱟᱹᱢᱤ ᱢᱮ; ᱟᱨᱠᱟᱭᱤᱵᱟᱞ ᱢᱩᱪᱟᱹᱫ ᱞᱟᱹᱜᱤᱫ ᱠᱚᱢᱟᱱᱰ ᱫᱚ `archive --expire 7y` ᱦᱩᱭ ᱫᱟᱲᱮᱭᱟᱜᱼᱟ ᱾
4. ᱢᱮᱴᱟᱣ ᱜᱷᱚᱴᱚᱱ ᱞᱚᱜᱽ ᱢᱮ, ᱡᱟᱦᱟᱸ ᱨᱮ ᱫᱟᱭᱤᱠ ᱚᱯᱟᱨᱮᱴᱚᱨ ID, ᱴᱟᱭᱤᱢᱥᱴᱮᱢᱯ, ᱟᱨ ᱥᱠᱳᱯ ᱥᱮᱞᱮᱫ ᱢᱮᱱᱟᱜᱼᱟ ᱾

ᱠᱟᱹᱢᱤ ᱱᱚᱴ ᱠᱚ:

- `PII` ᱟᱨ ᱮᱴᱟᱜ ᱥᱮᱱᱥᱤᱴᱤᱵᱽ ᱯᱷᱤᱞᱰ ᱠᱚ ᱞᱟᱹᱜᱤᱫ ᱢᱮᱴᱟᱣ ᱫᱚ ᱵᱟᱝ ᱨᱩᱣᱟᱹᱲ ᱜᱟᱱᱚᱜ ᱞᱮᱠᱟ ᱦᱩᱭᱩᱜᱼᱟ ᱾
- ᱰᱤᱥᱴᱨᱤᱵᱤᱭᱩᱴᱮᱰ ᱥᱤᱥᱴᱚᱢ ᱠᱚ ᱞᱟᱹᱜᱤᱫ, ᱯᱩᱥᱴᱟᱹᱣ ᱢᱮ ᱡᱮ ᱥᱟᱱᱟᱢ ᱨᱮᱯᱞᱤᱠᱟ ᱟᱨ ᱴᱟᱨᱥᱤᱭᱟᱨᱤ ᱥᱴᱳᱨᱮᱡᱽ ᱴᱤᱭᱟᱨ ᱠᱚ ᱢᱮᱴᱟᱣ ᱠᱟᱹᱢᱤ ᱨᱮ ᱥᱮᱞᱮᱫ ᱢᱮᱱᱟᱜᱼᱟ ᱾

## ᱵᱟᱫᱽ ᱟᱨ ᱚᱵᱷᱟᱨᱨᱟᱭᱤᱰ ᱠᱚ

- ᱟᱹᱱ, ᱨᱮᱜᱩᱞᱮᱴᱚᱨᱤ, ᱥᱮ ᱛᱚᱞᱟᱥ ᱦᱚᱞᱰ ᱠᱚ ᱚᱴᱚᱢᱮᱴᱮᱰ ᱢᱮᱴᱟᱣ ᱠᱚ ᱚᱵᱷᱟᱨᱨᱟᱭᱤᱰ ᱮᱫᱟ ᱾ ᱦᱚᱞᱰ ᱯᱨᱚᱠᱨᱤᱭᱟ ᱠᱚ ᱞᱟᱹᱜᱤᱫ ᱥᱮᱠᱥᱚᱱ 3.2 ᱧᱮᱞ ᱢᱮ ᱾
- ᱠᱷᱩᱞᱟᱹ ᱜᱨᱟᱦᱚᱠ ᱥᱮᱵᱟ ᱴᱤᱠᱮᱴ ᱠᱚ ᱞᱟᱹᱜᱤᱫ ᱞᱟᱹᱜᱛᱤᱭᱟᱱ ᱰᱮᱴᱟ ᱫᱚ ᱠᱮᱥ ᱪᱟᱵᱟᱜ ᱫᱷᱟᱹᱵᱤᱡ ᱫᱚᱦᱚ ᱜᱟᱱᱚᱜᱼᱟ, ᱛᱟᱭᱚᱢ ᱫᱚᱲᱦᱟ ᱧᱮᱞ ᱵᱤᱰᱟᱹᱣ ᱦᱩᱭᱩᱜᱼᱟ ᱾

## ᱵᱷᱩᱢᱤᱠᱟ ᱟᱨ ᱫᱟᱭᱤᱠ ᱠᱚ

- ᱰᱮᱴᱟ ᱢᱟᱹᱞᱤᱠ ᱠᱚ: ᱰᱮᱴᱟᱥᱮᱴ ᱠᱚ ᱞᱟᱹᱜᱤᱫ ᱨᱤᱴᱮᱱᱥᱚᱱ ᱠᱞᱟᱥ ᱠᱚ ᱮᱢ ᱢᱮ ᱟᱨ ᱵᱟᱫᱽ ᱠᱚ ᱢᱟᱱᱟᱣ ᱢᱮ ᱾
- IT ᱚᱯᱟᱨᱮᱥᱚᱱᱥ: ᱢᱮᱴᱟᱣ ᱠᱟᱹᱢᱤ ᱪᱟᱞᱟᱣ ᱢᱮ, ᱥᱩᱨᱩᱠᱷᱤᱛ ᱴᱩᱞᱤᱝ ᱫᱚᱦᱚᱭ ᱢᱮ, ᱟᱨ ᱯᱩᱨᱟᱹᱣ ᱯᱩᱥᱴᱟᱹᱣ ᱢᱮ ᱾
- ᱠᱚᱢᱯᱞᱟᱭᱮᱱᱥ ᱴᱤᱢ: ᱨᱤᱴᱮᱱᱥᱚᱱ ᱯᱨᱮᱠᱴᱤᱥ ᱠᱚ ᱚᱰᱤᱴ ᱢᱮ ᱟᱨ ᱚᱛᱷᱚᱨᱤᱴᱮᱴᱤᱵᱽ ᱨᱤᱴᱮᱱᱥᱚᱱ ᱥᱮᱰᱭᱩᱞ ᱫᱚᱦᱚᱭ ᱢᱮ ᱾

## ᱠᱚᱢᱯᱞᱟᱭᱮᱱᱥ, ᱚᱰᱤᱴ, ᱟᱨ ᱨᱤᱯᱳᱨᱴᱤᱝ

ᱚᱠᱛᱚ ᱚᱠᱛᱚ ᱨᱮ ᱚᱰᱤᱴ ᱠᱚ ᱨᱤᱴᱮᱱᱥᱚᱱ ᱠᱞᱟᱥ, ᱢᱮᱴᱟᱣ ᱞᱚᱜᱽ, ᱟᱨ ᱦᱚᱞᱰ ᱯᱨᱚᱠᱨᱤᱭᱟ ᱠᱚ ᱢᱟᱱᱟᱣ ᱵᱟᱛᱟᱣ ᱮᱫᱟ ᱾ ᱵᱟᱝ ᱢᱟᱱᱟᱣ ᱫᱚ ᱨᱤᱢᱮᱰᱤᱭᱮᱥᱚᱱ ᱯᱞᱟᱱ ᱠᱚ ᱮᱛᱚᱦᱚᱵᱟ ᱟᱨ ᱰᱤᱥᱤᱯᱞᱤᱱᱟᱨᱤ ᱠᱟᱹᱢᱤ ᱦᱚᱨᱟ ᱦᱩᱭ ᱫᱟᱲᱮᱭᱟᱜᱼᱟ ᱾

## FAQs

Q: ᱨᱤᱴᱮᱱᱥᱚᱱ ᱚᱠᱛᱚ ᱠᱚ ᱪᱮᱫ ᱞᱮᱠᱟ ᱴᱷᱟᱹᱣᱠᱟᱹᱜᱼᱟ?

A: ᱨᱤᱴᱮᱱᱥᱚᱱ ᱚᱠᱛᱚ ᱠᱚ ᱟᱹᱱ ᱞᱟᱹᱜᱛᱤ, ᱵᱮᱯᱟᱨ ᱞᱟᱹᱜᱛᱤ, ᱟᱨ ᱜᱚᱯᱚᱱᱤᱭᱚᱛᱟ ᱵᱤᱪᱟᱹᱨ ᱠᱚ ᱢᱤᱫ ᱥᱟᱶᱛᱮ ᱢᱮᱥᱟᱜᱼᱟ ᱾ ᱡᱚᱠᱷᱚᱱ ᱥᱚᱸᱫᱮᱦ ᱛᱟᱦᱮᱸᱱᱟ, ᱠᱚᱢᱯᱞᱟᱭᱮᱱᱥ ᱴᱤᱢ ᱥᱟᱶ ᱜᱟᱞᱢᱟᱨᱟᱣ ᱢᱮ ᱾

Q: ᱢᱮᱴᱟᱣ ᱛᱟᱭᱚᱢ ᱟᱨᱠᱟᱭᱤᱵᱽ ᱰᱮᱴᱟ ᱫᱚᱲᱦᱟ ᱧᱟᱢ ᱜᱟᱱᱚᱜᱼᱟ?

A: ᱵᱟᱝ ᱾ ᱢᱮᱴᱟᱣ ᱫᱚ ᱵᱟᱝ ᱨᱩᱣᱟᱹᱲ ᱜᱟᱱᱚᱜ ᱞᱮᱠᱟ ᱛᱮᱭᱟᱨ ᱟᱠᱟᱱᱟ ᱾ ᱟᱹᱰᱤ ᱞᱟᱹᱜᱛᱤᱭᱟᱱ ᱫᱚᱲᱦᱟ ᱧᱟᱢ ᱞᱟᱹᱜᱤᱫ, ᱟᱠᱚᱣᱟᱜ ᱨᱤᱴᱮᱱᱥᱚᱱ ᱣᱤᱱᱰᱳ ᱵᱷᱤᱛᱨᱤ ᱨᱮ ᱢᱮᱱᱟᱜ ᱵᱮᱠᱟᱯ ᱠᱚ ᱪᱮᱛᱟᱱ ᱨᱮ ᱵᱷᱚᱨᱥᱟᱭ ᱢᱮ ᱾

## ᱨᱤᱵᱷᱤᱡᱚᱱ ᱦᱤᱥᱴᱨᱤ

- ᱰᱨᱟᱯᱷᱴ 1.0: ᱵᱷᱤᱛᱨᱤ ᱧᱮᱞ ᱵᱤᱰᱟᱹᱣ ᱞᱟᱹᱜᱤᱫ ᱛᱮᱭᱟᱨ ᱟᱠᱟᱱ ᱯᱟᱹᱦᱤᱞ ᱱᱤᱛᱤ ᱱᱚᱴ ᱾
- ᱫᱚᱥᱟᱨ ᱨᱤᱵᱷᱤᱡᱚᱱ ᱨᱮ ᱵᱤᱥᱛᱟᱨ ᱨᱤᱴᱮᱱᱥᱚᱱ ᱢᱮᱴᱨᱤᱠᱥ ᱟᱨ ᱴᱮᱠᱱᱤᱠᱟᱞ ᱮᱯᱮᱱᱰᱤᱠᱥ ᱥᱮᱞᱮᱫ ᱛᱟᱦᱮᱸᱱᱟ ᱾

ᱵᱟᱹᱲᱛᱤ ᱫᱤᱥᱟᱹ ᱩᱫᱩᱜ ᱞᱟᱹᱜᱤᱫ, ᱰᱮᱴᱟ ᱜᱚᱵᱷᱟᱨᱱᱮᱱᱥ ᱚᱯᱷᱤᱥ ᱥᱟᱶ ᱡᱚᱜᱟᱡᱚᱜᱽ ᱢᱮ ᱥᱮ ᱪᱮᱛᱟᱱ ᱨᱮᱭᱟᱜ ᱞᱤᱝᱠ ᱨᱮ ᱯᱩᱨᱟᱹ ᱥᱮᱰᱭᱩᱞ ᱧᱮᱞ ᱢᱮ ᱾
English → Manipuri: Meetei Mayek script
# ꯗꯦꯇꯥ ꯔꯤꯇꯦꯟꯁꯟ ꯑꯃꯁꯨꯡ ꯗꯤꯂꯤꯁꯟ ꯄꯣꯂꯤꯁꯤ (ꯗ꯭ꯔꯥꯐꯠ)

## ꯁꯃꯔꯤ

ꯄꯣꯂꯤꯁꯤ ꯅꯣꯠ ꯑꯁꯤꯅ ꯑꯣꯔꯒꯥꯅꯥꯏꯖꯦꯁꯟꯒꯤ ꯗꯦꯇꯥ ꯔꯤꯇꯦꯟꯁꯟ, ꯄꯤꯔꯤꯑꯣꯗꯤꯛ ꯔꯤꯚꯤꯌꯨ, ꯑꯃꯁꯨꯡ ꯁꯦꯀ꯭ꯌꯨꯔ ꯗꯤꯂꯤꯁꯟꯒꯤ ꯑꯦꯞꯔꯣꯆ ꯑꯗꯨ ꯎꯠꯂꯤ꯫ ꯃꯁꯤꯅ ꯑꯦꯞꯂꯤꯀꯦꯕꯜ ꯂꯣꯁꯤꯡꯒ ꯀꯝꯄ꯭ꯂꯥꯏ ꯇꯧꯕ, ꯏꯟꯗꯤꯚꯤꯖꯨꯑꯦꯜ ꯄ꯭ꯔꯥꯏꯚꯦꯁꯤ ꯉꯥꯛꯊꯣꯛꯄ, ꯑꯃꯁꯨꯡ ꯍꯤꯁꯇꯣꯔꯤꯀꯦꯜ ꯔꯦꯀꯣꯔꯗꯁꯤꯡꯒꯤ ꯃꯊꯧ ꯇꯥꯗꯕ ꯁ꯭ꯇꯣꯔꯦꯖ ꯍꯟꯊꯍꯟꯅꯕ ꯄꯥꯟꯗꯝ ꯊꯝꯃꯤ꯫ ꯄꯣꯂꯤꯁꯤ ꯑꯁꯤ ꯑꯣꯔꯒꯥꯅꯥꯏꯖꯦꯁꯟꯅ ꯕꯤꯖꯤꯅꯦꯁꯀꯤ ꯃꯅꯨꯡꯗ ꯁꯦꯝꯈꯤꯕ, ꯐꯪꯈꯤꯕ, ꯅꯠꯇ꯭ꯔꯒ ꯃꯦꯟꯇꯦꯟ ꯇꯧꯈꯤꯕ ꯏꯂꯦꯛꯠꯔꯣꯅꯤꯛ ꯑꯃꯁꯨꯡ ꯐꯤꯖꯤꯀꯦꯜ ꯔꯦꯀꯣꯔꯗꯁꯤꯡꯗ ꯑꯦꯞꯂꯥꯏ ꯇꯧꯏ꯫

<a name="scope"></a>
## ꯁ꯭ꯀꯣꯞ

ꯄꯣꯂꯤꯁꯤ ꯑꯁꯤꯒꯤ ꯁ꯭ꯀꯣꯞꯇ ꯀꯁ꯭ꯇꯃꯔ ꯔꯦꯀꯣꯔꯗꯁꯤꯡ, ꯏꯝꯄ꯭ꯂꯣꯌꯤ ꯔꯦꯀꯣꯔꯗꯁꯤꯡ, ꯇ꯭ꯔꯥꯟꯖꯦꯛꯁꯅꯦꯜ ꯂꯣꯒꯁꯤꯡ, ꯕꯦꯀꯑꯞꯁꯤꯡ, ꯑꯃꯁꯨꯡ ꯑꯦꯅꯥꯂꯤꯇꯤꯛꯁ ꯗꯦꯇꯥ ꯌꯥꯎꯏ꯫ ꯃꯁꯤꯅ ꯁ꯭ꯄꯦꯁꯤꯑꯂꯥꯏꯖ ꯑꯥꯔꯀꯥꯏꯚ ꯑꯦꯒ꯭ꯔꯤꯃꯦꯟꯇꯁꯤꯡꯅ ꯀꯟꯠꯔꯣꯜ ꯇꯧꯕ ꯄꯦꯄꯔ ꯔꯦꯀꯣꯔꯗꯁꯤꯡ ꯌꯥꯎꯍꯟꯗꯦ, ꯃꯈꯣꯏꯕꯨ ꯁ꯭ꯄꯦꯁꯤꯐꯤꯀꯦꯜ ꯑꯣꯏꯅ ꯔꯤꯐꯔꯦꯟꯁ ꯇꯧꯗ꯭ꯔꯕꯗꯤ꯫ ꯑꯥꯏꯇꯦꯃꯥꯏꯖꯗ ꯁꯦꯗ꯭ꯌꯨꯜ ꯑꯃꯒꯤꯗꯃꯛ, [ꯔꯤꯇꯦꯟꯁꯟ ꯁꯦꯗ꯭ꯌꯨꯜ](https://company.example/policies/retention) ꯌꯦꯡꯕꯤꯌꯨ꯫

> ꯄ꯭ꯔꯤꯟꯁꯤꯄꯜ: ꯃꯊꯧ ꯇꯥꯕ ꯑꯗꯨꯈꯛꯇꯃꯛ ꯊꯝꯃꯨ, ꯃꯊꯧ ꯇꯥꯕ ꯃꯇꯝ ꯑꯗꯨ ꯐꯥꯎꯕ, ꯑꯃꯁꯨꯡ ꯃꯊꯧ ꯇꯥꯗꯕ ꯃꯇꯝꯗ ꯃꯁꯤꯕꯨ ꯁꯦꯀ꯭ꯌꯨꯔ ꯑꯣꯏꯅ ꯗꯤꯁꯄꯣꯖ ꯇꯧ꯫

## ꯗꯤꯐꯤꯅꯤꯁꯟꯁꯤꯡ

- `PII`: ꯃꯤꯡꯁꯤꯡ, ꯏꯃꯦꯜ ꯑꯦꯗ꯭ꯔꯦꯁꯁꯤꯡ, ꯅꯦꯁꯅꯦꯜ ꯑꯥꯏꯗꯦꯟꯇꯤꯐꯥꯏꯌꯔꯁꯤꯡ, ꯑꯃꯁꯨꯡ ꯃꯥꯟꯅꯕ ꯑꯥꯏꯗꯦꯟꯇꯤꯐꯥꯏꯌꯔꯁꯤꯡ ꯌꯥꯎꯅ ꯄꯔꯁꯣꯅꯦꯜ ꯑꯣꯏꯅ ꯑꯥꯏꯗꯦꯟꯇꯤꯐꯥꯏ ꯇꯧꯕ ꯌꯥꯕ ꯏꯟꯐꯣꯔꯃꯦꯁꯟ꯫
- ꯔꯤꯇꯦꯟꯁꯟ ꯄꯤꯔꯤꯑꯣꯗ: ꯗꯤꯂꯤꯁꯟ ꯇꯧꯗ꯭ꯔꯤꯉꯩꯒꯤ ꯃꯃꯥꯡꯗ ꯄꯤꯔꯤꯕ ꯔꯦꯀꯣꯔꯗ ꯑꯃ ꯊꯝꯒꯗꯕ ꯃꯊꯧ ꯇꯥꯕ ꯃꯇꯝ꯫
- ꯑꯥꯔꯀꯥꯏꯚꯦꯜ ꯁ꯭ꯇꯣꯔꯦꯖ: ꯔꯦꯀꯣꯔꯗꯁꯤꯡ ꯑꯦꯛꯇꯤꯚ ꯕꯤꯖꯤꯅꯦꯁ ꯌꯨꯖ ꯇꯧꯕꯒꯤ ꯃꯇꯨꯡꯗ ꯑꯗꯨꯕꯨ ꯐꯥꯏꯅꯦꯜ ꯗꯤꯂꯤꯁꯟꯒꯤ ꯃꯃꯥꯡꯗ ꯄ꯭ꯔꯤꯖꯔꯚ ꯇꯧꯕ ꯂꯣꯋꯔ-ꯀꯣꯁ꯭ꯠ ꯁ꯭ꯇꯣꯔꯦꯖ ꯇꯤꯌꯔ ꯑꯃꯅꯤ꯫

## ꯔꯤꯇꯦꯟꯁꯟ ꯁꯦꯗ꯭ꯌꯨꯜ (ꯍꯥꯏ ꯂꯦꯚꯦꯜ)

1. ꯀꯁ꯭ꯇꯃꯔ ꯑꯦꯀꯥꯎꯟꯠ ꯗꯦꯇꯥ: ꯑꯦꯀꯥꯎꯟꯠꯀꯤ ꯄꯨꯟꯁꯤ ꯑꯃꯁꯨꯡ ꯗꯤꯁ꯭ꯄ꯭ꯌꯨꯠ ꯔꯤꯖꯣꯂꯨꯁꯟꯒꯤꯗꯃꯛ ꯆꯍꯤ ꯲ ꯊꯝꯃꯨ꯫
2. ꯐꯥꯏꯅꯦꯟꯁꯤꯑꯦꯜ ꯑꯃꯁꯨꯡ ꯇꯦꯛꯁ ꯔꯦꯀꯣꯔꯗꯁꯤꯡ: ꯑꯦꯀꯥꯎꯟꯇꯤꯡ ꯔꯨꯜꯁꯤꯡꯒꯤ ꯃꯇꯨꯡ ꯏꯟꯅ ꯌꯥꯝꯗ꯭ꯔꯕꯗ ꯆꯍꯤ ꯷ ꯊꯝꯃꯨ꯫
3. ꯏꯝꯄ꯭ꯂꯣꯌꯤ ꯄꯔꯁꯣꯅꯦꯜ ꯔꯦꯀꯣꯔꯗꯁꯤꯡ: ꯂꯣ ꯑꯗꯨꯅ ꯑꯇꯣꯞꯄ ꯃꯊꯧ ꯇꯥꯗ꯭ꯔꯕꯗꯤ ꯇꯔꯃꯤꯅꯦꯁꯟꯒꯤ ꯃꯇꯨꯡꯗ ꯆꯍꯤ ꯶ ꯊꯝꯃꯨ꯫
4. ꯁꯤꯁ꯭ꯇꯦꯝ ꯂꯣꯒꯁꯤꯡ ꯑꯃꯁꯨꯡ ꯇꯦꯂꯤꯃꯦꯠꯔꯤ: ꯔꯣ ꯂꯣꯒꯁꯤꯡ ꯅꯨꯃꯤꯠ ꯹꯰ ꯊꯝꯃꯨ, ꯑꯦꯒ꯭ꯔꯤꯒꯦꯇꯦꯗ ꯇꯦꯂꯤꯃꯦꯠꯔꯤ ꯆꯍꯤ ꯳ ꯊꯝꯃꯨ꯫
5. ꯕꯦꯀꯑꯞꯁꯤꯡ: ꯔꯤꯇꯦꯟꯁꯟ ꯑꯁꯤ ꯕꯦꯀꯑꯞ ꯄꯣꯂꯤꯁꯤꯒꯤ ꯃꯇꯨꯡ ꯏꯟꯅ ꯆꯠꯂꯤ; ꯗꯤꯂꯤꯁꯟꯁꯤꯡ ꯑꯁꯤ ꯁꯦꯛꯁꯟ ꯳.꯲ ꯑꯃꯁꯨꯡ ꯗꯤꯁꯥꯁ꯭ꯇꯔ ꯔꯤꯀꯣꯚꯔꯤ ꯔꯤꯀ꯭ꯋꯥꯌꯔꯃꯦꯟꯇꯁꯤꯡꯅ ꯀꯟꯠꯔꯣꯜ ꯇꯧꯏ꯫

### ꯁ꯭ꯄꯦꯁꯤꯑꯦꯜ ꯀꯦꯇꯦꯒꯣꯔꯤꯁꯤꯡ

- HIPAA ꯒꯤ ꯃꯈꯥꯗ ꯂꯩꯕ ꯍꯦꯜꯊ ꯗꯦꯇꯥ: ꯂꯤꯒꯦꯜ ꯀꯥꯎꯟꯁꯦꯜ ꯒꯥꯏꯗꯦꯟꯁ ꯑꯃꯁꯨꯡ ꯁꯦꯀ꯭ꯌꯨꯔ ꯗꯤꯂꯤꯁꯟ ꯄ꯭ꯔꯣꯇꯣꯀꯣꯜꯁꯤꯡꯒꯤ ꯃꯇꯨꯡ ꯏꯟꯅ ꯊꯝꯃꯨ ꯑꯃꯁꯨꯡ ꯗꯤꯂꯤꯠ ꯇꯧ꯫
- ꯂꯤꯇꯤꯒꯦꯁꯟ ꯍꯣꯜꯗꯒꯤ ꯃꯈꯥꯗ ꯂꯩꯕ ꯗꯦꯇꯥ: ꯍꯣꯜꯗ ꯑꯗꯨ ꯂꯤꯐꯠ ꯇꯧꯗ꯭ꯔꯤꯐꯥꯎꯕ ꯔꯦꯒꯨꯂꯔ ꯗꯤꯂꯤꯁꯟ ꯁꯁꯄꯦꯟꯗ ꯇꯧ꯫

## ꯗꯦꯇꯥ ꯗꯤꯂꯤꯁꯟ ꯄ꯭ꯔꯣꯁꯤꯖꯔꯁꯤꯡ

<a name="deletion-procedure"></a>
ꯔꯦꯀꯣꯔꯗ ꯑꯃꯅ ꯃꯁꯤꯒꯤ ꯔꯤꯇꯦꯟꯁꯟ ꯄꯤꯔꯤꯑꯣꯗꯀꯤ ꯑꯔꯣꯏꯕꯗ ꯌꯧꯔꯛꯄ ꯃꯇꯝꯗ, ꯃꯁꯤꯕꯨ ꯃꯈꯥꯒꯤ ꯁ꯭ꯇꯦꯞꯁꯤꯡ ꯑꯁꯤ ꯁꯤꯖꯤꯟꯅꯗꯨꯅ ꯗꯤꯂꯤꯠ ꯇꯧꯒꯗꯕꯅꯤ ꯅꯠꯇ꯭ꯔꯒ ꯑꯦꯅꯣꯅꯤꯃꯥꯏꯖ ꯇꯧꯒꯗꯕꯅꯤ:

1. ꯔꯤꯇꯦꯟꯁꯟ ꯃꯦꯇꯥꯗꯦꯇꯥ ꯚꯦꯔꯤꯐꯥꯏ ꯇꯧ ꯑꯃꯁꯨꯡ ꯍꯣꯜꯗꯁꯤꯡ ꯅꯠꯇ꯭ꯔꯒ ꯑꯦꯛꯁꯦꯞꯁꯟꯁꯤꯡ ꯑꯃꯠꯇ ꯑꯦꯞꯂꯥꯏ ꯇꯧꯗꯦ ꯍꯥꯏꯕ ꯀꯟꯐꯔꯝ ꯇꯧ꯫
2. ꯏꯂꯤꯖꯤꯕꯜ ꯑꯣꯏꯔꯕꯗꯤ, ꯔꯦꯀꯣꯔꯗ ꯑꯗꯨ ꯁꯦꯀ꯭ꯌꯨꯔ ꯗꯤꯂꯤꯁꯟ ꯀ꯭ꯌꯨꯗ ꯍꯣꯡꯗꯣꯛꯎ꯫
3. ꯑꯦꯞꯔꯨꯚ ꯇꯧꯔꯕ ꯇꯨꯜꯁꯤꯡ ꯁꯤꯖꯤꯟꯅꯗꯨꯅ ꯗꯤꯂꯤꯁꯟ ꯑꯦꯛꯁꯤꯀ꯭ꯌꯨꯠ ꯇꯧ; ꯑꯥꯔꯀꯥꯏꯚꯦꯜ ꯑꯦꯛꯁꯄꯥꯏꯔꯤꯒꯤꯗꯃꯛ ꯈꯨꯗꯝ ꯑꯣꯏꯕ ꯀꯃꯥꯟꯗ ꯑꯁꯤ `archive --expire 7y` ꯑꯣꯏꯕ ꯌꯥꯏ꯫
4. ꯔꯦꯁꯄꯣꯟꯁꯤꯕꯜ ꯑꯣꯄꯔꯦꯇꯔ ID, ꯇꯥꯏꯝꯁ꯭ꯇꯦꯝꯄ, ꯑꯃꯁꯨꯡ ꯁ꯭ꯀꯣꯞ ꯌꯥꯎꯅ ꯗꯤꯂꯤꯁꯟ ꯏꯚꯦꯟꯠ ꯑꯗꯨ ꯂꯣꯒ ꯇꯧ꯫

ꯑꯣꯄꯔꯦꯁꯅꯦꯜ ꯅꯣꯠꯁꯤꯡ:

- `PII` ꯑꯃꯁꯨꯡ ꯑꯇꯣꯞꯄ ꯁꯦꯟꯁꯤꯇꯤꯚ ꯐꯤꯜꯗꯁꯤꯡꯒꯤꯗꯃꯛ ꯗꯤꯂꯤꯁꯟ ꯑꯁꯤ ꯔꯤꯚꯔꯁꯤꯕꯜ ꯑꯣꯏꯗꯕ ꯑꯣꯏꯒꯗꯕꯅꯤ꯫
- ꯗꯤꯁ꯭ꯠꯔꯤꯕ꯭ꯌꯨꯇꯦꯗ ꯁꯤꯁ꯭ꯇꯦꯝꯁꯤꯡꯒꯤꯗꯃꯛ, ꯔꯦꯄ꯭ꯂꯤꯀꯥꯁꯤꯡ ꯑꯃꯁꯨꯡ ꯇꯔꯁꯤꯌꯔꯤ ꯁ꯭ꯇꯣꯔꯦꯖ ꯇꯤꯌꯔꯁꯤꯡ ꯄꯨꯝꯅꯃꯛ ꯗꯤꯂꯤꯁꯟ ꯋꯥꯔꯛꯐꯀꯤ ꯃꯅꯨꯡꯗ ꯌꯥꯎꯍꯟꯕ ꯁꯣꯌꯗꯅ ꯇꯧ꯫

## ꯑꯦꯛꯁꯦꯞꯁꯟꯁꯤꯡ ꯑꯃꯁꯨꯡ ꯑꯣꯚꯔꯔꯥꯏꯗꯁꯤꯡ

- ꯂꯤꯒꯦꯜ, ꯔꯦꯒꯨꯂꯦꯇꯔꯤ, ꯅꯠꯇ꯭ꯔꯒ ꯏꯟꯚꯦꯁꯇꯤꯒꯦꯇꯔꯤ ꯍꯣꯜꯗꯁꯤꯡꯅ ꯑꯣꯇꯣꯃꯦꯇꯦꯗ ꯗꯤꯂꯤꯁꯟꯕꯨ ꯑꯣꯚꯔꯔꯥꯏꯗ ꯇꯧꯏ꯫ ꯍꯣꯜꯗ ꯄ꯭ꯔꯣꯁꯤꯖꯔꯁꯤꯡꯒꯤꯗꯃꯛ ꯁꯦꯛꯁꯟ ꯳.꯲ ꯌꯦꯡꯕꯤꯌꯨ꯫
- ꯑꯣꯄꯟ ꯀꯁ꯭ꯇꯃꯔ ꯁꯔꯚꯤꯁ ꯇꯤꯀꯦꯠꯁꯤꯡꯒꯤꯗꯃꯛ ꯃꯊꯧ ꯇꯥꯕ ꯗꯦꯇꯥ ꯑꯁꯤ ꯀꯦꯁ ꯀ꯭ꯂꯣꯖꯔ ꯇꯧꯗ꯭ꯔꯤꯐꯥꯎꯕ ꯊꯝꯕ ꯌꯥꯏ, ꯃꯗꯨꯒꯤ ꯃꯇꯨꯡꯗ ꯑꯃꯨꯛ ꯍꯟꯅ ꯏꯚꯦꯜꯌꯨꯑꯦꯠ ꯇꯧꯒꯅꯤ꯫

## ꯔꯣꯜꯁꯤꯡ ꯑꯃꯁꯨꯡ ꯔꯦꯁꯄꯣꯟꯁꯤꯕꯤꯂꯤꯇꯤꯁꯤꯡ

- ꯗꯦꯇꯥ ꯑꯣꯅꯔꯁꯤꯡ: ꯗꯦꯇꯥꯁꯦꯠꯁꯤꯡꯗ ꯔꯤꯇꯦꯟꯁꯟ ꯀ꯭ꯂꯥꯁꯁꯤꯡ ꯑꯦꯁꯥꯏꯟ ꯇꯧ ꯑꯃꯁꯨꯡ ꯑꯦꯛꯁꯦꯞꯁꯟꯁꯤꯡ ꯑꯦꯞꯔꯨꯚ ꯇꯧ꯫
- IT ꯑꯣꯄꯔꯦꯁꯟꯁꯤꯡ: ꯗꯤꯂꯤꯁꯟ ꯋꯥꯔꯛꯐꯀꯁꯤꯡ ꯔꯟ ꯇꯧ, ꯁꯦꯀ꯭ꯌꯨꯔ ꯇꯨꯂꯤꯡ ꯃꯦꯟꯇꯦꯟ ꯇꯧ, ꯑꯃꯁꯨꯡ ꯀꯝꯄ꯭ꯂꯤꯠꯅꯦꯁ ꯚꯦꯔꯤꯐꯥꯏ ꯇꯧ꯫
- ꯀꯝꯄ꯭ꯂꯥꯏꯑꯦꯟꯁ ꯇꯤꯝ: ꯔꯤꯇꯦꯟꯁꯟ ꯄ꯭ꯔꯦꯛꯇꯤꯁꯁꯤꯡ ꯑꯣꯗꯤꯠ ꯇꯧ ꯑꯃꯁꯨꯡ ꯑꯣꯊꯣꯔꯤꯇꯦꯇꯤꯚ ꯔꯤꯇꯦꯟꯁꯟ ꯁꯦꯗ꯭ꯌꯨꯜ ꯃꯦꯟꯇꯦꯟ ꯇꯧ꯫

## ꯀꯝꯄ꯭ꯂꯥꯏꯑꯦꯟꯁ, ꯑꯣꯗꯤꯠ, ꯑꯃꯁꯨꯡ ꯔꯤꯄꯣꯔꯇꯤꯡ

ꯄꯤꯔꯤꯑꯣꯗꯤꯛ ꯑꯣꯗꯤꯠꯁꯤꯡꯅ ꯔꯤꯇꯦꯟꯁꯟ ꯀ꯭ꯂꯥꯁꯁꯤꯡ, ꯗꯤꯂꯤꯁꯟ ꯂꯣꯒꯁꯤꯡ, ꯑꯃꯁꯨꯡ ꯍꯣꯜꯗ ꯄ꯭ꯔꯣꯁꯤꯖꯔꯁꯤꯡꯒ ꯀꯝꯄ꯭ꯂꯥꯏ ꯇꯧꯕ ꯑꯗꯨ ꯚꯦꯂꯤꯗꯦꯠ ꯇꯧꯒꯅꯤ꯫ ꯀꯝꯄ꯭ꯂꯥꯏ ꯇꯧꯗꯕꯅ ꯔꯤꯃꯦꯗꯤꯑꯦꯁꯟ ꯄ꯭ꯂꯥꯟꯁꯤꯡ ꯇ꯭ꯔꯤꯒꯔ ꯇꯧꯒꯅꯤ ꯑꯃꯁꯨꯡ ꯗꯤꯁꯤꯄ꯭ꯂꯤꯅꯔꯤ ꯑꯦꯛꯁꯟ ꯊꯣꯛꯍꯟꯕ ꯌꯥꯏ꯫

## FAQs

Q: ꯔꯤꯇꯦꯟꯁꯟ ꯄꯤꯔꯤꯑꯣꯗꯁꯤꯡ ꯀꯔꯝꯅ ꯂꯦꯞꯄꯤ?

A: ꯔꯤꯇꯦꯟꯁꯟ ꯄꯤꯔꯤꯑꯣꯗꯁꯤꯡꯅ ꯂꯤꯒꯦꯜ ꯔꯤꯀ꯭ꯋꯥꯌꯔꯃꯦꯟꯇꯁꯤꯡ, ꯕꯤꯖꯤꯅꯦꯁ ꯃꯊꯧ ꯇꯥꯕ, ꯑꯃꯁꯨꯡ ꯄ꯭ꯔꯥꯏꯚꯦꯁꯤ ꯀꯟꯁꯤꯗꯔꯦꯁꯟꯁꯤꯡ ꯄꯨꯟꯁꯤꯟꯅꯩ꯫ ꯆꯤꯡꯅꯕ ꯂꯩꯔꯕꯗꯤ, ꯀꯝꯄ꯭ꯂꯥꯏꯑꯦꯟꯁ ꯇꯤꯝꯒ ꯇꯥꯟꯅꯕꯤꯌꯨ꯫

Q: ꯑꯥꯔꯀꯥꯏꯚ ꯇꯧꯔꯕ ꯗꯦꯇꯥ ꯑꯁꯤ ꯗꯤꯂꯤꯁꯟ ꯇꯧꯔꯕ ꯃꯇꯨꯡꯗ ꯔꯤꯁ꯭ꯇꯣꯔ ꯇꯧꯕ ꯌꯥꯕ꯭ꯔꯥ?

A: ꯌꯥꯗꯦ꯫ ꯗꯤꯂꯤꯁꯟ ꯑꯁꯤ ꯔꯤꯚꯔꯁꯤꯕꯜ ꯑꯣꯏꯗꯕ ꯑꯣꯏꯅ ꯄꯥꯟꯗꯝ ꯊꯝꯃꯤ꯫ ꯀ꯭ꯔꯤꯇꯤꯀꯦꯜ ꯔꯤꯁ꯭ꯇꯣꯔꯦꯁꯟꯁꯤꯡꯒꯤꯗꯃꯛ, ꯃꯈꯣꯏꯒꯤ ꯔꯤꯇꯦꯟꯁꯟ ꯋꯤꯟꯗꯣꯒꯤ ꯃꯅꯨꯡꯗ ꯂꯩꯕ ꯕꯦꯀꯑꯞꯁꯤꯡꯗ ꯊꯥꯖꯕ ꯊꯝꯃꯨ꯫

## ꯔꯤꯚꯤꯖꯟ ꯍꯤꯁ꯭ꯇꯔꯤ

- ꯗ꯭ꯔꯥꯐꯠ ꯱.꯰: ꯏꯟꯇꯔꯅꯦꯜ ꯔꯤꯚꯤꯌꯨꯒꯤꯗꯃꯛ ꯁꯦꯝ-ꯁꯥꯔꯕ ꯑꯍꯥꯟꯕ ꯄꯣꯂꯤꯁꯤ ꯅꯣꯠ꯫
- ꯃꯊꯪꯒꯤ ꯔꯤꯚꯤꯖꯟꯗ ꯗꯤꯇꯦꯜ ꯔꯤꯇꯦꯟꯁꯟ ꯃꯦꯠꯔꯤꯛꯁ ꯑꯃꯁꯨꯡ ꯇꯦꯛꯅꯤꯀꯦꯜ ꯑꯦꯄꯦꯟꯗꯤꯁꯤꯡ ꯌꯥꯎꯒꯅꯤ꯫

ꯑꯍꯦꯟꯕ ꯒꯥꯏꯗꯦꯟꯁꯀꯤꯗꯃꯛ, ꯗꯦꯇꯥ ꯒꯣꯚꯔꯅꯦꯟꯁ ꯑꯣꯐꯤꯁꯇ ꯀꯟꯇꯦꯛ ꯇꯧꯕꯤꯌꯨ ꯅꯠꯇ꯭ꯔꯒ ꯃꯊꯛꯀꯤ ꯂꯤꯡꯛ ꯑꯗꯨꯗ ꯂꯩꯕ ꯃꯄꯨꯡ ꯐꯥꯕ ꯁꯦꯗ꯭ꯌꯨꯜ ꯑꯗꯨ ꯌꯦꯡꯕꯤꯌꯨ꯫

All structural elements are preserved in all six outputs: the ten ## headings, one ### heading, fourteen bullets, both numbered lists, both anchors, the Markdown link, the bare URL, and the inline code spans PII and archive --expire 7y.

This is true for both low-resource and high-resource languages. Section references (Section 3.2) and retention periods (7 years, 90 days) are also unchanged.


Prompting guide

The model is designed to work with a simple, consistent prompting format. Following these guidelines ensures the best translation quality.

1. Specify only the target language

Use the following prompt format:

Translate the following text into {target_language}:

{source_text}

Do not mention the source language. For example, prefer:

Translate the following text into Hindi:

Hello world.

instead of:

Translate the following text from English into Hindi:

Hello world.

This one format covers every capability the model supports, only the target-language name and the verb change:

Translation (native script, either direction):

Translate the following text into Hindi:

Hello world.
Translate the following text into English:

नमस्ते दुनिया

Romanized translation: append (Roman) to the target language name:

Translate the following text into Hindi (Roman):

Hello world.

Romanized input needs no marker of its own; ask for the target you want (native script or (Roman), or English) the same way as above.

Transliteration: same (Roman) convention, with "Transliterate" instead of "Translate." This changes script only, not language or meaning:

Transliterate the following text into Hindi (Roman):

नमस्ते दुनिया
Transliterate the following text into Hindi:

namaste duniya

Code-mixed translation: append (code-mixed) to the target language name to get output that mixes English and the target language within a sentence, the way people commonly type when switching languages or without an Indic keyboard. Code-mixed input needs no marker; translating it into a native-script target (or English) works with the plain format above:

Translate the following text into Hindi (code-mixed):

Can you send me the report by tomorrow evening?
Translate the following text into English:

Yaar tomorrow ka meeting cancel ho gaya hai.

See Supported languages for the full set of qualified target names.

2. Use a single user message

Pass the translation request as a single user turn, without a system message.

messages = [
    {
        "role": "user",
        "content": "Translate the following text into Hindi:\n\nHello world."
    }
]

Adding a system message, even an empty one, changes the rendered prompt and can reduce translation quality.

3. Use the provided chat template

Render the prompt using the packaged chat_template.jinja with add_generation_prompt=True. The rendered prompt has the following structure:

<bos><|turn>user
Translate the following text into Hindi:

Hello world.<turn|>
<|turn>model

The model generates the translation directly and terminates with <turn|>. Both <turn|> and <eos> are configured as end-of-sequence tokens (eos_token_id = [1, 106]). Optionally including <turn|> as an explicit stop string is also supported.

The model produces only the translated text, it does not generate reasoning, intermediate steps, or explanatory content.


Environment setup

One script builds a clean conda environment with everything the model needs:

bash envs/setup_inference_env.sh      # creates conda env `indic-translate-inference`

It needs transformers 5.12 or newer (the release that ships the Gemma 4 architecture), torch 2.6 or newer, and Python 3.10 or newer. The script picks the CUDA build that matches your driver.

For anything beyond this, installing by hand, driver and CUDA troubleshooting, Docker, and checking an existing environment, see the bodhan_gen_ai_tools repository.


Basic inference

import torch
from transformers import AutoModelForImageTextToText, AutoProcessor

MODEL = "bodhan-ai/indic-translate"

processor = AutoProcessor.from_pretrained(MODEL)
model = AutoModelForImageTextToText.from_pretrained(
    MODEL,
    dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="sdpa",
).eval()


def translate(text: str, target_language: str, max_new_tokens: int = 512) -> str:
    messages = [{
        "role": "user",
        "content": f"Translate the following text into {target_language}:\n\n{text}",
    }]
    inputs = processor.apply_chat_template(
        messages,
        add_generation_prompt=True,
        tokenize=True,
        return_dict=True,
        return_tensors="pt",
    ).to(model.device)

    with torch.inference_mode():
        output = model.generate(
            **inputs,
            max_new_tokens=max_new_tokens,
            do_sample=False,        # greedy, recommended for translation
            use_cache=True,
        )

    completion = output[0, inputs["input_ids"].shape[-1]:]
    return processor.decode(completion, skip_special_tokens=True).strip()


print(translate("The committee approved the proposal after a long debate.", "Hindi"))
print(translate("समिति ने लंबी बहस के बाद प्रस्ताव को मंजूरी दे दी।", "English"))

That is all that is needed to translate one piece of text. Pass the target language by name, the full list is under Supported languages, and do not mention the source language; the model works it out.

For high-throughput serving with vLLM, batch inference, and the ready-made command-line scripts, see the bodhan_gen_ai_tools repository.


Workload temperature top_p max_tokens max_model_len
Sentence / short segment 0.0 (greedy) 1.0 512 8,192
Paragraph 0.0 (greedy) 1.0 2,048 16,384
Full document 0.00.2 1.0 8,192 32,768

Greedy decoding is the default recommendation: it is the strongest setting for translation, and the only one that reproduces exactly run to run.

Set max_tokens generously for Indic targets. Most Indic scripts tokenise less efficiently than English, so a translation typically requires 1.5–2× the token count of its English source. Leave enough room and the whole translation comes back.


Hardware requirements

Weights (bfloat16) 15.9 GB
Minimum GPU 24 GB (e.g. L4, RTX 4090): sentence and paragraph workloads
Recommended GPU 40 GB or more (e.g. A100 40/80 GB, H100): full 32k-token documents
Multi-GPU Supported via vLLM --tensor-parallel-size; not required
Precision bfloat16, unquantised

A single 40 GB card serves full 32k-token documents comfortably, and one 24 GB card is enough for sentence and paragraph workloads, so the model fits on hardware most teams already have.


Performance

Measured 2026-08-30 on one NVIDIA H100 80GB HBM3, vLLM 0.20.2, bfloat16, max_model_len=32768, CUDA graphs enabled, greedy decoding. Requests are real translation traffic, IN22-Gen segments for the sentence workload, whole ~4,000-character documents for the document workload, spread across eight target languages so the output-length mix is not best-case. Latency is measured at the client with streaming enabled; throughput is measured non-streaming, which is how batch pipelines run.

The raw measurements are in benchmarks/: every figure below can be traced to the run that produced it, including the full spread of scores. Reproduce any row with scripts/benchmark_serving.py; see benchmarks/README.md for the method.

Reading the tables. TTFT is the wait before the first word appears. E2E is the total time for the whole translation. p50 is the middle of the measurements, half the requests were faster than that and half slower. p95 is near the slow end: 95% of requests were faster than that number, so it shows what the unluckiest requests looked like. Both are given because a good middle can hide a slow tail.

Interactive latency

One request at a time, streaming: the numbers a user feels:

Time to first token Full response Decode rate
Sentence (~50 tokens in) 16 ms 0.26 s 177 tok/s
Document (~980 tokens in, ~1,490 out) 35 ms 7.50 s 165 tok/s

First token comes back in well under a twentieth of a second, so a translation starts appearing effectively as soon as it is requested. A complete 4,000-character document, around 600 words of formatted Markdown or LaTeX, lands in about seven and a half seconds.

Sentence latency under load

Streaming. 4,096 distinct requests per level at concurrency 1–64, 20,000 at concurrency 128–512. Streaming is the right choice for interactive surfaces, where the first token matters more than peak tokens per second:

Concurrency Output tok/s Sentences/s TTFT p50 TTFT p95 E2E p50 E2E p95
1 170 3.4 16 ms 17 ms 0.26 s 0.62 s
8 1,268 24.9 24 ms 26 ms 0.28 s 0.66 s
32 4,306 85.3 35 ms 81 ms 0.33 s 0.76 s
64 6,104 121.2 143 ms 247 ms 0.47 s 0.91 s
128 6,744 132.3 358 ms 562 ms 0.94 s 1.34 s
256 7,204 141.4 595 ms 999 ms 1.81 s 2.29 s
512 7,696 151.3 1,059 ms 1,797 ms 3.38 s 4.26 s

Time to first token stays under 150 ms at up to 64 concurrent requests, and a complete sentence returns in 0.47 s at that level, so a single GPU sustains around 120 interactive translations per second before latency starts climbing noticeably.

Sentence throughput

Non-streaming, 8,192 distinct requests per level, ~50 output tokens each. Disabling per-token streaming events substantially increases throughput, which is the appropriate configuration for bulk pipelines:

Concurrency Output tok/s Sentences/s E2E p50 E2E p95
128 12,299 242.0 0.45 s 1.05 s
256 14,442 283.9 0.68 s 1.94 s
512 14,320 281.6 1.35 s 4.39 s

Over 1.0 million sentences per hour on a single GPU at concurrency 256 (the peak of this sweep, concurrency 512 is not faster once the engine is already saturated), with median latency still under a second. The engine's own accounting peaked at 18,870 output tok/s (~28,500 tokens/s including prefill) across this run.

Document throughput

Whole documents per request, ~1,490 output tokens each (non-streaming):

Concurrency Output tok/s Documents/s E2E p50 E2E p95
16 2,107 1.60 9.30 s 11.76 s
64 4,710 3.64 11.45 s 14.77 s

Roughly 13,100 documents per hour on one GPU, and per-document latency grows modestly between 16 and 64 concurrent requests as the batch absorbs the load.

Concurrency headroom

At 32k context, vLLM allocates 47.6 GiB of KV cache = 1,873,565 tokens, enough for 57 simultaneous full-length 32,768-token documents on a single card. Typical document and sentence traffic uses a small fraction of that: KV-cache occupancy stayed under 7% throughout the sweeps above, so the limit in practice is compute, not memory.


Limitations

These should be reviewed before designing an evaluation; several affect what to test and how to interpret the results.

  • Performance depends on the documented prompt format, and degrades silently. This is the most common avoidable cause of poor output. A non-conforming prompt still returns fluent, plausible text, no error is raised and there is nothing to catch. The only symptom is quality quietly below what the model can deliver. Import scripts/indic_translate_prompt.py and build requests with it rather than hand-writing prompt strings.

  • It is a translation model, not an assistant. It takes one instruction and one piece of source text and returns the translation. It does not answer questions about the text, summarise it, or follow multi-step instructions, and prompting it to do so gives unreliable results.

  • Quality varies by language. The widely-resourced languages are markedly stronger. The lower-resource ones, Kashmiri, Sanskrit, Santali, Sindhi, Manipuri, trail the rest in both directions and are worth checking more closely.

  • Translation into English is stronger than translation out of it. Allow more review effort for English → Indic than for the reverse direction.

  • Code and docstrings can come back partly untranslated. In documents that mix prose with code, the translation itself is generally sound, but fragments inside the code, docstrings, inline comments, short string literals, are sometimes returned in English rather than the target language. No content is corrupted or omitted; the affected fragments are returned unchanged. Review code-heavy documents for residual English before accepting the output.

  • Romanized document translation occasionally fails to complete. At full 32k-token document length, 1.7–6.0% of generations in the Romanized translation/transliteration directions were truncated or malformed, versus effectively none in the native-script document workload. Sentence-level Romanized requests were unaffected. See Romanized translation & transliteration quality.

  • Digit rendering is not uniform across languages. Some target languages receive native script digits and others Latin digits, and Kannada produces both. Any downstream system that parses numbers, dates or quantities out of the translation should normalise digits first. See Example translations for the full breakdown.


License and citation

Released under Indic Open Model License v1.0.

The base model is Gemma-4-E4B-it. You should confirm that your use of this model conforms the terms of use and license of these upstream models/components also.


If you find the license difficult to understand, here is a plain-language guide to the Indic Open Model License.

Broad, no-cost access for research, government, nonprofit, and commercial use — with a few conditions attached.

This deed is a human-readable summary of the license, not a substitute for it. Where the two disagree, the full Indic Open Model License governs.


You're free to

No cost, no royalty, worldwide — for research, government, nonprofit, and commercial use, at any scale.

  • Run it — for inference, in a product, in research, however you like.
  • Change it — fine-tune, distill, quantize, merge, or otherwise build on it.
  • Self-host it — power your own product or service with it, commercial or not.
  • Share it — pass on copies of the model or your own version of it.

As long as you

Five conditions cover almost everything. The rest of the license is these, spelled out in legal detail.

1. Give credit

Wherever you ship the model or a derivative to anyone else, say where it came from — and don't strip out existing notices.

"Built with [Model Name] from Bodhan AI / AI4Bharat."

2. Pass it on the same way

If you give your fine-tuned or derived version to anyone else — hand it over, or run it as a service for them — it carries this exact license. You can't relicense it on different terms.

3. Ask before hosting it for others

Self-hosting is free. But if you're going to run it as an API or hosted service that other people or companies call directly, that needs Bodhan AI's written sign-off first — unless you're a nonprofit, government, or academic user, or you publicly release an equally capable open version within 90 days.

4. Don't use it to cause harm

No exceptions — not even for nonprofit or research use. That means no:

  • child sexual abuse material, or content that sexualizes minors
  • weapons development, including chemical, biological, radiological, or nuclear
  • mass surveillance or social-scoring systems
  • disinformation campaigns, including election manipulation
  • automated decisions that affect someone's legal rights without human oversight
  • deepfakes or voice clones of real people without their consent
  • robocalls, auto-dialers, or voice-phishing scams
  • AI companion products designed to simulate romance or foster emotional dependency

5. Talk to us if your product gets huge

If your own product built on this — not through hosting it for others, that's covered above — crosses either threshold, you'll need a separate commercial license. Doesn't apply to nonprofit, government, or academic users.

Threshold
500M+ monthly active users
or
$250M+ annual revenue

Citation

@misc{indic-translate-2026,
  title  = {Indic-Translate: Document-Level Machine Translation for 22 Indian Languages},
  author = {Bodhan AI and Ai4Bharat},
  year   = {2026},
  url    = {https://bodhan.ai/research/blogs/indic-translate}
}

Support

For access requests, evaluation questions, or quality reports, contact the Bodhan AI language technology team. When reporting a translation issue, please include the source text, the target language, the exact prompt sent, your decoding parameters, and the output you received, the prompt format is the most common source of unexpected quality loss.


Downloads last month
76
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bodhan-ai/indic-translate

Finetuned
(349)
this model

Collection including bodhan-ai/indic-translate