Model Card for gemma-2-2b-jpn-it-translate-gguf

gemma-2-2b-jpn-it-translate-ggufใฏใ€ๆ—ฅ่‹ฑใƒป่‹ฑๆ—ฅ็ฟป่จณใ‚ฟใ‚นใ‚ฏใซ็‰นๅŒ–ใ—ใŸSLM๏ผˆSmall Language Model๏ผ‰ใงใ™ใ€‚ใƒ‘ใƒฉใƒกใƒผใ‚ฟใƒผๆ•ฐใฏ20ๅ„„๏ผˆ2B๏ผ‰ใงใ™ใŒใ€ๅˆ†้‡Žใซใ‚ˆใฃใฆใฏๅพ“ๆฅใฎ70ๅ„„๏ผˆ7B๏ผ‰ใƒขใƒ‡ใƒซใซ่ฟซใ‚‹ใƒฌใƒ™ใƒซใฎ็ฟป่จณๅ“่ณชใ‚’ๆไพ›ใ—ใพใ™ใ€‚ใƒ•ใ‚กใ‚คใƒซใ‚ตใ‚คใ‚บใŒ็ด„2GB็จ‹ๅบฆใงใ‚ใ‚‹ใŸใ‚ๆฏ”่ผƒ็š„ๅฐใ•ใ„ใŸใ‚ใ€้ซ˜้€ŸใชๅฎŸ่กŒใŒๅฏ่ƒฝใงใ™ใ€‚
gemma-2-2b-jpn-it-translate-gguf is an SLM (Small Language Model) specialized for Japanese-English and English-Japanese translation tasks. Despite having only 2 billion parameters (2B), it provides translation quality approaching that of conventional 7 billion (7B) parameter models in some kind of text. With a relatively small file size of about 2GB, it enables fast execution.

ๆ–‡ๅ˜ไฝใง็ฟป่จณใ™ใ‚‹ไบ‹ใ‚’ๅญฆ็ฟ’ใ—ใฆใ„ใ‚‹ใŸใ‚ใ€ๆ”น่กŒใ‚’ๅซใ‚€้•ทๆ–‡ใ‚’ไธ€ๅบฆใซๆธกใ™ใจๅ“่ณชใŒไฝŽไธ‹ใ—ใพใ™ใ€‚
้•ทๆ–‡ใ‚’็ฟป่จณใ™ใ‚‹้š›ใฏๆ–‡ๅ˜ไฝใงๅŒบๅˆ‡ใ‚‹ๅ‰ๅ‡ฆ็†ใ‚’ใ—ใฆใ‹ใ‚‰ใƒขใƒ‡ใƒซใซไธŽใˆใฆใใ ใ•ใ„ใ€‚
Because the model is trained to translate sentence by sentence, passing a long sentence with line breaks at once will result in a decrease in quality.
When translating a long sentence, please pre-process it by dividing it into sentences before feeding it to the model.

Sample Colab Script

Google ใ‚ขใ‚ซใ‚ฆใƒณใƒˆใ‚’ใŠๆŒใกใฎๆ–นใฏไปฅไธ‹ใฎใƒชใƒณใ‚ฏๅ…ˆใงOpen in Colabใƒœใ‚ฟใƒณใ‚’ๆŠผใ™ไบ‹ใง่ฉฆใ™ไบ‹ใŒใงใใพใ™
If you have a Google account, you can try it out by clicking the Open in Colab button at the link below.
Colab sample

sample for windows

ColabใฎCPUใฏ้…ใ„ใฎใงใ”่‡ช่บซใฎใƒ‘ใ‚ฝใ‚ณใƒณใงllama.cppใ‚’ใ‚ณใƒณใƒ‘ใ‚คใƒซใ—ใฆๅ‹•ใ‹ใ™ๆ–นใŒๅฟซ้ฉใงใ™ใ€‚
Colab's CPU is slow, so it is more convenient to compile and run llama.cpp on your own computer.

ใ‚ฏใƒฉใ‚คใ‚ขใƒณใƒˆ๏ผใ‚ตใƒผใƒใƒผๅฝขๆ…‹ใงๅ‹•ไฝœใ•ใ›ใ‚‹ใ‚ตใƒณใƒ—ใƒซใฏไปฅไธ‹ใงใ™ใ€‚
Below is a sample of how it works in client/server mode.

start server.

.\llama.cpp\build\bin\Release\llama-server -m .\gemma-2-2b-jpn-it-translate-Q4_K_L.gguf -c 2048 --override-kv tokenizer.ggml.add_bos_token=bool:false

bosใƒˆใƒผใ‚ฏใƒณ้‡่ค‡ใ‚’ๅ›ž้ฟใ™ใ‚‹ใŸใ‚ใซใ€--override-kv tokenizer.ggml.add_bos_token=bool:false ใ‚ชใƒ—ใ‚ทใƒงใƒณใ‚’ๅฟ…ใšๆŒ‡ๅฎšใ—ใฆใใ ใ•ใ„ใ€‚
Be sure --override-kv tokenizer.ggml.add_bos_token=bool:false options for avoid dup bos token.

pip install -U transformers
pip install requests
import transformers
import requests
import json
from transformers import AutoTokenizer

system_prompt = "You are a highly skilled professional Japanese-English and English-Japanese translator. Translate the given text accurately, taking into account the context and specific instructions provided. Steps may include hints enclosed in square brackets [] with the key and value separated by a colon:. Only when the subject is specified in the Japanese sentence, the subject will be added when translating into English. If no additional instructions or context are provided, use your expertise to consider what the most appropriate context is and provide a natural translation that aligns with that context. When translating, strive to faithfully reflect the meaning and tone of the original text, pay attention to cultural nuances and differences in language usage, and ensure that the translation is grammatically correct and easy to read. After completing the translation, review it once more to check for errors or unnatural expressions. For technical terms and proper nouns, either leave them in the original language or use appropriate translations as necessary. Take a deep breath, calm down, and start translating.\n\n"
instruct = "Translate Japanese to English.\nWhen translating, please use the following hints:\n[writing_style: casual]"

initial_messages  = [
    {"role": "user", "content": system_prompt + instruct},
    {"role": "assistant", "content": "OK"}
]

message_list = [
 "and I was a little bit nervous, too, speaking to a Japanese audience really for the first time, certainly since I left the White House. ",
 "And I had a very good interpreter, and if you have ever made a speech in Japan in English, it takes a lot longer to say it in Japanese. ",
 "I decided I would break the ice by telling the shortest joke that I knew.",
 "It was not the best joke I knew, but it was the shortest joke I knew, left over from my governor's campaign years before.",
 "So I told my joke, the interpreter told the joke, and the audience just collapsed in laughter. ",
 "I never got a better response from any audience in my life. ",
 "So I could not wait to get through the speech and talk to the interpreter and ask him,"
 "\"How did you tell my joke?\"",
 "He was very evasive. He would not tell me how he told it.",
 "I insisted, and he finally ducked his head and said,",
 "\"I told the audience, 'President Carter told a funny story. Everybody, laugh.'\""
]

tokenizer = AutoTokenizer.from_pretrained("webbigdata/gemma-2-2b-jpn-it-translate")

if __name__ == "__main__":
    messages = initial_messages.copy()
    for i in range(len(message_list)):
        messages.append({"role": "user",  "content": message_list[i]})
        print("user: " + message_list[i])


        # Transformersใฎใƒˆใƒผใ‚ฏใƒŠใ‚คใ‚ถใƒผใ‚’ไฝฟใ„ใŸใใชใ„ๅ ดๅˆใฏใ€Colabใฎใ‚ตใƒณใƒ—ใƒซใŒๆ‰‹ใงใƒ—ใƒญใƒณใƒ—ใƒˆใƒ†ใƒณใƒ—ใƒฌใƒผใƒˆใ‚’ๆ›ธใ„ใฆใ‚‹ใฎใงใใกใ‚‰ใ‚’ๅ‚่€ƒใ—ใฆใใ ใ•ใ„
        # If you donโ€™t want to use the Transformers tokenizer, you can use the Colab example to manually write the prompt template.  
        prompt = tokenizer.apply_chat_template(
            messages,
            add_generation_prompt=True,
            tokenize=False
        )

        payload = {
            "prompt": prompt,
            "n_predict": 1200
        }
    
        # Define the URL and headers for the POST request
        url = "http://localhost:8080/completion"
        headers = {
            "Content-Type": "application/json"
        }

        # Send the POST request and capture the response
        response = requests.post(url, headers=headers, data=json.dumps(payload))
        # print(response)
        # print( response.json() )


        # Check if the request was successful
        if response.status_code != 200:
            print(f"Error: {response.text}")

        # Parse the response JSON
        response_data = response.json()

        # Extract the 'content' field from the response
        response_content = response_data.get('content', '').strip()

        print("assistant: " + response_content)
        messages.append({"role": "assistant",  "content": response_content})
        
        # Max 6 message, you need more memory for more massages.
        if len(messages) > 8:  # 2 (initial) + 6 (new) = 8
            messages = initial_messages + messages[-6:]

result

user: and I was a little bit nervous, too, speaking to a Japanese audience really for the first time, certainly since I left the White House.
assistant: ใใ—ใฆใ€็งใ‚‚ใกใ‚‡ใฃใจ็ทŠๅผตใ—ใฆใ„ใพใ—ใŸใ€‚็™ฝๅฎฎใ‚’้›ขใ‚Œใฆใ‹ใ‚‰ใ€ๆ—ฅๆœฌไบบใฎๅ‰ใงๅˆใ‚ใฆใ€ๆœฌๅฝ“ใฎๆ„ๅ‘ณใง่ฉฑใ—ใฆใ„ใ‚‹ใจใ„ใ†ใฎใฏ ใ€ใ‹ใชใ‚Šๅˆใงใ—ใŸใ€‚
user: And I had a very good interpreter, and if you have ever made a speech in Japan in English, it takes a lot longer to say it in Japanese.
assistant: ใใ—ใฆใ€็งใฏใจใฆใ‚‚ๅ„ชใ‚ŒใŸ้€š่จณใŒใ„ใฆใ€ใ‚‚ใ—ใ‚ใชใŸใŒ่‹ฑ่ชžใงๆ—ฅๆœฌใง่ฌ›ๆผ”ใ‚’ใ—ใŸใ‚Šใ—ใŸใ“ใจใŒใ‚ใ‚‹ใฎใชใ‚‰ใ€ๆ—ฅๆœฌ่ชžใง่ฉฑใ™ใฎ ใซใฏใ‚‹ใ‹ใซ้•ทใใชใ‚Šใพใ™ใ€‚
user: I decided I would break the ice by telling the shortest joke that I knew.
assistant: ใใ“ใงใ€ๆœ€ใ‚‚็Ÿญใ„ใ‚ธใƒงใƒผใ‚ฏใ‚’่ฉฑใ—ใฆใ€ๆœ€ๅˆใฎใ†ใกใฏๆฐ—ใ‚’็ทฉใพใ›ใ‚ˆใ†ใจๆฑบใ‚ใŸใฎใงใ™ใ€‚
user: It was not the best joke I knew, but it was the shortest joke I knew, left over from my governor's campaign years before.
assistant: ไธ€็•ชใฎใ‚ธใƒงใƒผใ‚ฏใงใฏใชใ„ใงใ™ใŒใ€็งใฎ็Ÿฅใฃใฆใ„ใŸ็Ÿญใ„ใ‚ธใƒงใƒผใ‚ฏใ€ใใ‚Œใพใง governors ใฎ้ธๆŒ™ใงใฎใ‚ญใƒฃใƒณใƒšใƒผใƒณไธญใซใ„ใŸใ‚‚ใฎใงใ™ใ€‚
user: So I told my joke, the interpreter told the joke, and the audience just collapsed in laughter.
assistant: ใใ‚Œใงใ€็งใŒ็งใฎใ‚ธใƒงใƒผใ‚ฏใ‚’่จ€ใฃใŸใ‚‰ใ€้€š่จณใŒ็งใฎใ‚ธใƒงใƒผใ‚ฏใ‚’ใ€ใใ—ใฆใ‚ชใƒผใƒ‡ใ‚ฃใ‚จใƒณใ‚นใฏใŸใ ใŸใ ็ฌ‘ใ„่ปขใ’ใพใ—ใŸใ€‚
user: I never got a better response from any audience in my life.
assistant: ็งใฏใ“ใฎใ‚ธใƒงใƒผใ‚ฏใฏใ€ไปŠใพใงใซใ€ใ“ใ‚ŒไปฅไธŠใ„ใ„ๅๅฟœใ‚’ใ‚‚ใ‚‰ใฃใŸใ“ใจใŒใ‚ใ‚Šใพใ›ใ‚“ใงใ—ใŸใ€‚
user: So I could not wait to get through the speech and talk to the interpreter and ask him,"How did you tell my joke?"
assistant: ใใ‚Œใงใ€็งใฏใ€ใใฎๆผ”่ชฌใ‚’็ต‚ใ‚ใ‚‰ใ›ใฆใ€้€š่จณใจ่ฉฑใ™ใฎใŒๅพ…ใกใใ‚Œใชใ„ใงใ€ใใฎ้€š่จณใซๅฐ‹ใญใŸใ„ใจๆ€ใฃใฆใ„ใพใ—ใŸใ€ใ€Œใฉใฎ ใ‚ˆใ†ใซ็งใฎใ‚ธใƒงใƒผใ‚ฏใ‚’ไผใˆใ‚‹ใฎ๏ผŸใ€
user: He was very evasive. He would not tell me how he told it.
assistant: ๅฝผใฏใจใฆใ‚‚้ฟใ‘ใฆใ„ใพใ—ใŸใ€‚ใฉใ†ใ—ใŸใฎใ‹ใจๆ•™ใˆใฆใใ‚Œใพใ›ใ‚“ใ€‚
user: I insisted, and he finally ducked his head and said,
assistant: ๆ–ญใ‚‹ใพใ„ใจใ€ๅฝผใฏใคใ„ใซ้ฆ–ใ‚’ใ‹ใ—ใ’ใฆใ€ใ€Œ
user: "I told the audience, 'President Carter told a funny story. Everybody, laugh.'"
assistant: ใ€Œๅธไผš่€…ใ‹ใ‚‰ใ€ใ‚ซใƒผใ‚ฟใƒผๅคง็ตฑ้ ˜ใŒ้ข็™ฝใ„่ฉฑใ‚’ใ—ใŸใ€‚็š†ใ•ใ‚“ใ€็ฌ‘ใฃใฆใ€‚ใ€

ใƒ™ใƒณใƒใƒžใƒผใ‚ฏ็ตๆžœ Benchmark results

Q4KL

filename direction spBLEU chrF2++ comet xlcomet
flores200v1 enja 21.94 30.8 0.8714 0.7496
flores200v1 jaen 21.28 52.2 0.8577 0.8965
wmt23 jaen 14.46 41.5 0.7859 0.8353
wmt20 enja 14.15 25.2 0.8519 0.6701
wmt23 enja 13.96 25.4 0.8339 0.7500
Business jaen 19.14 42.7 0.8084 0.8309
wmt20 jaen 13.23 41.8 0.7807 0.7187
Business enja 17.53 33.7 0.8809 0.8604

่ฌ่พž Acknowledgements

BibTeX:

@misc{dahara2024imatrix,
  author       = {dahara1@webbigdata},
  title        = {gemma-2-2b-jpn-it-translate: A translation task-specific gguf model based on gemma-2-2b-jpn-it},
  year         = {2024},
  howpublished = {\url{https://huggingface.co/webbigdata/gemma-2-2b-jpn-it-translate-gguf/}},
  note         = {Accessed: 2024-10-10},
  abstract     = {This model was developed to verify how much Japanese-English and English-Japanese translation performance can be improved with the 2B gguf model.},
}
Downloads last month
541
GGUF
Model size
3B params
Architecture
gemma2
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for webbigdata/gemma-2-2b-jpn-it-translate-gguf

Quantized
(23)
this model