TRIEM_AI / docs /VIVA_PRESENTATION_SCRIPT.md
rzvn1's picture
Testing CodeRabbit
6f18524
|
Raw
History Blame Contribute Delete
6.33 kB

🎀 TRIEM AI - Viva Presentation Script & Code Walkthrough

Use this script during your presentation. It is structured to help you explain the code line-by-line or block-by-block to the examiner.


🟒 Introduction (Opening Statement)

"Good morning/afternoon. My project is TRIEM AI (Tribal Responsive Intelligent Empowerment Model). It is a Voice-to-Voice AI Assistant specifically designed for the Santali language. Unlike generic assistants, TRIEM focuses on:

  1. Local Language Support: Understanding Santali using AI4Bharat models.
  2. Hybrid Intelligence: Working both Online (Groq/Gemini) and Offline (Ollama/Local Cache).
  3. Empowerment: Helping tribal communities access information through simple voice interaction."

πŸ—οΈ Architecture Overview (The "Body" Analogy)

"Before I show the code, I want to explain how the system works. I designed it like a human body:

  • The Ears (asr_provider.py): Listens to Santali audio.
  • The Brain (brain.py): Thinks and generates answers.
  • The Mouth (tts_provider.py): Speaks the answer back.
  • The Translator (mt_provider.py): Translates Santali to English (for thinking) and back.
  • The Heart (server.py): Connects everything together."

πŸ’» Code Walkthrough (Detailed Explanation)

(Open server.py on your screen)

1. server.py - The Main Controller

"This is the main entry point. It's a Flask web server."

Key Lines to Point At:

app = Flask(__name__, ...)

"Here, I initialize the web application. static_folder serves my CSS/JS, and template_folder serves the HTML."

def init_models():

"This function runs when the app starts. It loads our heavy AI models into memory once so we don't reload them for every request. This makes the app faster."

@app.route('/api/process', methods=['POST'])

"This is the most important endpoint. It receives the user's audio file from the frontend."

def process_audio():

"Let me explain the pipeline inside this function step-by-step:

  1. Data Collection: I get request.files['audio'].
  2. ASR: I call asr_model.transcribe() to get the text from audio.
  3. Translation: I call mt_model.translate() to convert Santali to English.
  4. Intelligence: I send the English text to get_ai_response() (The Brain).
  5. Response: The AI gives an answer.
  6. Translation Back: I translate the English answer back to Santali.
  7. TTS: Finally, tts_model.speak_to_file() converts that text into an audio file for the user to hear."

(Open src/brain.py)

2. src/brain.py - The Intelligence Engine

"This file decides what to say. It handles the logic for Online vs. Offline modes."

Key Lines to Point At:

faq_manager = FAQManager()

"I initialize a local database manager here. This is my Cache Memory."

def get_ai_response(text, ...):

"This function takes the user's query and finds the best answer."

(Line ~86) if santali_text: ... faq_manager.get_answer()

"First, I check if we already have the answer in our local database. If yes, I return it immediately. This is zero-latency and works offline!"

(Line ~128) use_groq = mode in ["auto", "online"] ...

"If we are online, I use Groq. It's an ultra-fast AI inference engine."

(Line ~161) if mode in ["auto", "offline"]: ...

"If internet fails or if the user selected 'Offline Mode', I switch to Ollama. This runs a localized Llama-3 model directly on this laptop, ensuring privacy and reliability."


(Open src/asr_provider.py)

3. src/asr_provider.py - The ASR Module

"This module handles speech recognition using the IndicConformer model."

Key Lines to Point At:

self.model = AutoModel.from_pretrained(Config.ASR_MODEL_NAME ...)

"I am using the Hugging Face transformers library to load the model defined in my config file."

wav, sr = torchaudio.load(audio_file_path)

"I use TorchAudio to load the .wav file. I also resample it to 16kHz because that is what the AI model requires."

transcription = self.model(wav, language='sat')

"This is the actual inference. I explicitly pass language='sat' to tell the model to listen for Santali specifically."


(Open src/mt_provider.py)

4. src/mt_provider.py - The Translator

"Since most large AI models understand English best, translation is crucial."

Key Lines to Point At:

self.ip = IndicProcessor(inference=True)

"I use IndicTransToolkit. It handles the complex script issues often found in Indian languages like Santali (Ol Chiki script)."

batch = self.ip.preprocess_batch( ... src_lang='sat_Olck', tgt_lang='eng_Latn')

"Here, I prepare the text. notice the codes: sat_Olck stands for Santali in Ol Chiki script. eng_Latn is English."

model.generate(**inputs ...)

"The model then predicts the translated sentence."


(Open src/tts_provider.py)

5. src/tts_provider.py - The Text-to-Speech

"This gives the AI a voice."

Key Lines to Point At:

description = "A female speaker delivers a very slow ... in Santali."

"I am using ParlerTTS. It is a unique model that takes a text description of the voice. I programmed it to speak slowly and clearly so tribal users can understand easily."

sf.write(output_file, audio_arr ...)

"Finally, I save the generated sound wave as a .wav file, which is sent back to the user's browser."


(Open web/static/script.js)

6. The Frontend Logic (script.js)

"This JavaScript file makes the website interactive."

Key Functions:

navigator.mediaDevices.getUserMedia({ audio: true })

"This line asks the browser for permission to use the microphone."

mediaRecorder.ondataavailable

"As the user speaks, we collect the audio data into chunks."

drawVisualizer()

"This function draws the real-time audio wave animation you see on screen using the HTML5 Canvas API."


🏁 Conclusion

"To summarize:

  1. Frontend captures audio.
  2. Server orchestrates the process.
  3. ASR/MT/TTS models handle the language barriers.
  4. Brain provides the intelligence via Hybrid Cloud/Local AI.

This architecture ensures TRIEM is Robust, Scalable, and Culturally Responsive. Thank you!"