File size: 4,735 Bytes
ee292f1
 
1140d7d
ee292f1
 
 
53f745d
ee292f1
 
 
 
0643685
 
41c120f
0643685
 
 
26d3d03
0643685
 
 
 
 
 
26d3d03
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9546896
26d3d03
 
 
 
0643685
 
 
 
55f4cda
 
 
 
 
 
 
 
 
 
 
07185a6
 
55f4cda
 
0643685
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
26d3d03
0643685
 
 
 
 
 
 
26d3d03
0643685
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
---
title: "Linguini"
emoji: "🔊"
colorFrom: "purple"
colorTo: "gray"
sdk: "gradio"
sdk_version: "4.37.2"
app_file: stream_app.py
pinned: false
---

# Audio Stream Project

This project is a Gradio-based application that allows users to interact with an AI-powered audio streaming service. It utilizes OpenAI's language models for various tasks, such as speech recognition, text generation, language understanding, and speech generation.

## Table of Contents

- [Install](#install)
- [Usage](#usage)
- [Features](#features)
- [Configuration](#configuration)
- [Deployment](#deployment)
- [License](#license)

## Install
Installation and environment setup is currently done locally, but with a little effort, we could 
make a docker image.

1. Install miniconda or anaconda [here](https://docs.conda.io/en/latest/miniconda.html)
2. Create a new environment with the following command:
    ```bash
    conda create -n audio-stream python=3.11
    ```
3. Activate the environment:
    ```bash
    conda activate audio-stream
    ```
4. Install the required packages:
    ```bash
    pip install -r requirements.txt
    ```
5. Set the required environment variables or create a `.env` file (exclude the word `export` for `.env`)
    ```bash
    export OPENAI_API_KEY=<your_openai_api_key>
    ```
6. Run the Gradio application:
    ```bash
    python stream_app.py
    ```

### General Requirements
- Python 3.9 or higher  (recommend 3.11)
- [Gradio](https://www.gradio.app/) library
- [OpenAI](https://openai.com/) API key
- [Whisper](https://github.com/openai/whisper) library (for speech recognition)

### Running local Ollama
Ollama is great for running local models that are tuned or high performance versus those that are running online.  Generally, there is a three step process of getting ollama running, downloading the right llm model to use, and locally launching the model with litellm.  The library litellm provides the glue between ollama and the a programmatic interface for you to access locally.

1. Download [ollama](https://ollama.com/) 
2. Find the model you want to use and install it via the commandn line ``ollama pull <model>``
3. Run it locally with the command ``ollama serve``
  * If you run the local application (a lamma appears in your menu/run items), you may not need to explicitly run the serve command.
  * After launching, you can confirm that ollama is running on at this endpoint `127.0.0.1:11434`

As of 5/25/24, some models to consider are [lamma3](https://ollama.com/library/llama3) for general conversations and [dolphin-llama3](https://ollama.com/library/dolphin-llama3) for coding tasks.  Runner up mentions are [Microsoft's wizard2](https://ollama.com/library/wizardlm2) and [llava-llama3](https://ollama.com/library/llava-llama3)

### Microphone access error
Your browser might prevent you from access the microphone when running locally (http). To update this in chrome: Update chrome flags (chrome://flags) and allow local (http://127.0.0.1:7860) to be treated as secure (Insecure origins treated as secure)



## Usage

1. Set the required environment variables:
- `OPENAI_API_KEY`: Your OpenAI API key.
2. Run the Gradio application:
3. Access the application in your web browser at the provided URL.

## Features

- **Speech-to-Text Transcription**: Users can upload or record audio files, and the application will transcribe the speech using the Whisper model.
- **Text-to-Speech Generation**: Users can input text, and the application will generate audio output using OpenAI's language models.
- **Language Understanding**: The application can analyze the user's input and provide insights, such as sentiment analysis, topic detection, and entity recognition.
- **Conversational AI**: Users can engage in interactive conversations with the AI assistant, which can respond to queries, provide information, and assist with various tasks.

## Configuration

The application can be configured using command-line arguments or environment variables. Run the main command with the option `--help` to get a full list of available options.

## Deployment

The application can be deployed to various platforms, such as:

- **Local Machine**: Run the application on your local machine using the instructions in the [Usage](#usage) section.
- **Docker**: Package the application in a Docker container for easy deployment and scaling.
- **Cloud Platform**: Deploy the application to a cloud platform like AWS, Google Cloud, or Azure.  At first blush [a deployment strategy like this one](https://vinaykachare.medium.com/serverless-api-with-aws-sam-fastapi-3f4d9510d6b6) seems like a good follow-up for automated deployment.


## License

This project is licensed under the [MIT License](LICENSE).