Instructions to use patdev/k3-a40-bootstrap with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use patdev/k3-a40-bootstrap with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: llama cli -hf patdev/k3-a40-bootstrap:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./llama-cli -hf patdev/k3-a40-bootstrap:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf patdev/k3-a40-bootstrap:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf patdev/k3-a40-bootstrap:BF16
Use Docker
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- LM Studio
- Jan
- Ollama
How to use patdev/k3-a40-bootstrap with Ollama:
ollama run hf.co/patdev/k3-a40-bootstrap:BF16
- Unsloth Desktop
- Docker Model Runner
How to use patdev/k3-a40-bootstrap with Docker Model Runner:
docker model run hf.co/patdev/k3-a40-bootstrap:BF16
- Lemonade
How to use patdev/k3-a40-bootstrap with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull patdev/k3-a40-bootstrap:BF16
Run and chat with the model
lemonade run user.k3-a40-bootstrap-BF16
List all available models
lemonade list
- Atomic Chat
journal de demarrage
Browse files- etat/iquduwt33wwpvy.log +49 -1
etat/iquduwt33wwpvy.log
CHANGED
|
@@ -1,4 +1,4 @@
|
|
| 1 |
-
=== bootstrap v84-nom-nemotron | 11:22:
|
| 2 |
[VL] 11:14:42 bootstrap v84-nom-nemotron
|
| 3 |
[VL] 11:14:42 pilote 595.71.05, CUDA runtime 13.2
|
| 4 |
[VL] 11:14:42 pilote 595.71.05 : compat CUDA non necessaire
|
|
@@ -16,6 +16,7 @@
|
|
| 16 |
[VL] 11:15:51 cache de compilation, cle 94035774d7d7763f (NVIDIA_RTX_PRO_5000_Blackwell, vllm 0.27.1)
|
| 17 |
[VL] 11:15:51 pas de cache pour cette cle : premiere compilation, il sera publie ensuite
|
| 18 |
[VL] 11:15:51 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/iquduwt33wwpvy.log
|
|
|
|
| 19 |
|
| 20 |
=== nvidia-smi ===
|
| 21 |
41310 MiB, 48935 MiB
|
|
@@ -126,5 +127,52 @@
|
|
| 126 |
(APIServer pid=5630) INFO 08-30 11:22:08 [api_server.py:678] Supported tasks: ['generate']
|
| 127 |
(APIServer pid=5630) INFO 08-30 11:22:08 [parser_manager.py:37] "auto" tool choice has been enabled.
|
| 128 |
(APIServer pid=5630) INFO 08-30 11:22:15 [hf.py:540] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 129 |
|
| 130 |
=== proxy.log (fin) ===
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
=== bootstrap v84-nom-nemotron | 11:22:54 UTC ===
|
| 2 |
[VL] 11:14:42 bootstrap v84-nom-nemotron
|
| 3 |
[VL] 11:14:42 pilote 595.71.05, CUDA runtime 13.2
|
| 4 |
[VL] 11:14:42 pilote 595.71.05 : compat CUDA non necessaire
|
|
|
|
| 16 |
[VL] 11:15:51 cache de compilation, cle 94035774d7d7763f (NVIDIA_RTX_PRO_5000_Blackwell, vllm 0.27.1)
|
| 17 |
[VL] 11:15:51 pas de cache pour cette cle : premiere compilation, il sera publie ensuite
|
| 18 |
[VL] 11:15:51 journal distant : https://huggingface.co/patdev/k3-a40-bootstrap/resolve/main/etat/iquduwt33wwpvy.log
|
| 19 |
+
[VL] 11:22:52 READY qwen38nvfp4
|
| 20 |
|
| 21 |
=== nvidia-smi ===
|
| 22 |
41310 MiB, 48935 MiB
|
|
|
|
| 127 |
(APIServer pid=5630) INFO 08-30 11:22:08 [api_server.py:678] Supported tasks: ['generate']
|
| 128 |
(APIServer pid=5630) INFO 08-30 11:22:08 [parser_manager.py:37] "auto" tool choice has been enabled.
|
| 129 |
(APIServer pid=5630) INFO 08-30 11:22:15 [hf.py:540] Detected the chat template content format to be 'openai'. You can set `--chat-template-content-format` to override this.
|
| 130 |
+
(APIServer pid=5630) INFO 08-30 11:22:44 [base.py:235] Multi-modal warmup completed in 29.022s
|
| 131 |
+
(APIServer pid=5630) INFO 08-30 11:22:45 [base.py:235] Readonly multi-modal warmup completed in 0.736s
|
| 132 |
+
(APIServer pid=5630) WARNING 08-30 11:22:45 [model.py:1637] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 1.0, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.
|
| 133 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [api_server.py:682] Starting vLLM server on http://0.0.0.0:18081
|
| 134 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:37] Available routes are:
|
| 135 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /openapi.json, Methods: HEAD, GET
|
| 136 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /docs, Methods: HEAD, GET
|
| 137 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /docs/oauth2-redirect, Methods: HEAD, GET
|
| 138 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /redoc, Methods: HEAD, GET
|
| 139 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /load, Methods: GET
|
| 140 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /version, Methods: GET
|
| 141 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /health, Methods: GET
|
| 142 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /metrics, Methods: GET
|
| 143 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /tokenize, Methods: POST
|
| 144 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /detokenize, Methods: POST
|
| 145 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/models, Methods: GET
|
| 146 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /ping, Methods: GET
|
| 147 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /ping, Methods: POST
|
| 148 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /invocations, Methods: POST
|
| 149 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/chat/completions, Methods: POST
|
| 150 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/chat/completions/batch, Methods: POST
|
| 151 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/responses, Methods: POST
|
| 152 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/responses/{response_id}, Methods: GET
|
| 153 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/responses/{response_id}/cancel, Methods: POST
|
| 154 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/completions, Methods: POST
|
| 155 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/messages, Methods: POST
|
| 156 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/messages/count_tokens, Methods: POST
|
| 157 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /generative_scoring, Methods: POST
|
| 158 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /scale_elastic_ep, Methods: POST
|
| 159 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /is_scaling_elastic_ep, Methods: POST
|
| 160 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/chat/completions/render, Methods: POST
|
| 161 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/completions/render, Methods: POST
|
| 162 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/chat/completions/derender, Methods: POST
|
| 163 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /v1/completions/derender, Methods: POST
|
| 164 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:46] Route: /inference/v1/generate, Methods: POST
|
| 165 |
+
(APIServer pid=5630) INFO 08-30 11:22:46 [launcher.py:99] API server: waiting for HTTP server to start
|
| 166 |
+
(APIServer pid=5630) INFO: Started server process [5630]
|
| 167 |
+
(APIServer pid=5630) INFO: Waiting for application startup.
|
| 168 |
+
(APIServer pid=5630) INFO: Application startup complete.
|
| 169 |
+
(APIServer pid=5630) INFO 08-30 11:22:47 [launcher.py:105] API server: HTTP server started
|
| 170 |
+
(APIServer pid=5630) INFO: 127.0.0.1:58586 - "GET /health HTTP/1.1" 200 OK
|
| 171 |
+
(APIServer pid=5630) INFO: 127.0.0.1:46254 - "GET /health HTTP/1.1" 200 OK
|
| 172 |
+
(APIServer pid=5630) INFO: 127.0.0.1:46258 - "GET /health HTTP/1.1" 200 OK
|
| 173 |
|
| 174 |
=== proxy.log (fin) ===
|
| 175 |
+
INFO: Started server process [7771]
|
| 176 |
+
INFO: Waiting for application startup.
|
| 177 |
+
INFO: Application startup complete.
|
| 178 |
+
INFO: Uvicorn running on http://0.0.0.0:8088 (Press CTRL+C to quit)
|