MiniMax Agent commited on
Commit
41831f1
·
1 Parent(s): d36a46f

Add complete local Ollama setup with OpenELM - includes setup script, API server, test scripts, and documentation

Browse files
Dockerfile.api ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Dockerfile for the API server that uses Ollama
2
+
3
+ FROM python:3.10-slim
4
+
5
+ WORKDIR /app
6
+
7
+ # Install dependencies
8
+ COPY requirements_local.txt .
9
+ RUN pip install --no-cache-dir -r requirements_local.txt
10
+
11
+ # Copy application
12
+ COPY app_ollama.py .
13
+
14
+ # Expose port
15
+ EXPOSE 8000
16
+
17
+ # Run the API server
18
+ CMD ["python", "app_ollama.py"]
README_LOCAL.md ADDED
@@ -0,0 +1,393 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Complete Ollama OpenELM Setup Guide
2
+
3
+ This guide provides complete instructions to set up a local Ollama instance with Apple's OpenELM model and use it via OpenAI/Anthropic compatible APIs.
4
+
5
+ ## Table of Contents
6
+
7
+ 1. [Prerequisites](#prerequisites)
8
+ 2. [Quick Start (One Command)](#quick-start-one-command)
9
+ 3. [Manual Setup](#manual-setup)
10
+ 4. [Testing](#testing)
11
+ 5. [API Usage](#api-usage)
12
+ 6. [Docker Compose Setup](#docker-compose-setup)
13
+ 7. [Troubleshooting](#troubleshooting)
14
+
15
+ ---
16
+
17
+ ## Prerequisites
18
+
19
+ ### Required Software
20
+
21
+ - **Docker**: [Install Docker](https://docs.docker.com/get-docker/)
22
+ - **NVIDIA Driver**: For GPU support (check with `nvidia-smi`)
23
+ - **NVIDIA Container Toolkit**: [Install Guide](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/install-guide.html)
24
+
25
+ ### Verify GPU Access
26
+
27
+ ```bash
28
+ # Check NVIDIA driver
29
+ nvidia-smi
30
+
31
+ # Verify Docker can see GPU
32
+ docker run --rm --gpus all nvidia/cuda:11.0-base nvidia-smi
33
+ ```
34
+
35
+ ---
36
+
37
+ ## Quick Start (One Command)
38
+
39
+ ### For Linux/macOS
40
+
41
+ ```bash
42
+ # Download and run the complete setup script
43
+ curl -O https://raw.githubusercontent.com/your-repo/setup_ollama_openelm.sh
44
+ chmod +x setup_ollama_openelm.sh
45
+ ./setup_ollama_openelm.sh
46
+ ```
47
+
48
+ ### For Windows (PowerShell)
49
+
50
+ ```powershell
51
+ # Run each command manually (see Manual Setup below)
52
+ ```
53
+
54
+ ---
55
+
56
+ ## Manual Setup
57
+
58
+ ### Step 1: Start Ollama Container
59
+
60
+ ```bash
61
+ # Start Ollama with GPU support
62
+ docker run -d \
63
+ --name ollama \
64
+ -v ollama:/root/.ollama \
65
+ -p 127.0.0.1:11434:11434 \
66
+ --gpus all \
67
+ ollama/ollama
68
+
69
+ # Verify it's running
70
+ docker ps | grep ollama
71
+ ```
72
+
73
+ ### Step 2: Pull OpenELM Model
74
+
75
+ ```bash
76
+ # Pull the 3B parameter model (2.1 GB)
77
+ docker exec -it ollama ollama pull apple/OpenELM-3B-Instruct
78
+
79
+ # Verify installation
80
+ docker exec ollama ollama list
81
+ ```
82
+
83
+ Expected output:
84
+ ```
85
+ NAME ID SIZE MODIFIED
86
+ apple/OpenELM-3B-Instruct:latest abc123... 2.1 GB About a minute ago
87
+ ```
88
+
89
+ ### Step 3: Install Python Dependencies
90
+
91
+ ```bash
92
+ # Create virtual environment (optional but recommended)
93
+ python3 -m venv ollama_env
94
+ source ollama_env/bin/activate # Linux/macOS
95
+ # or: .\ollama_env\Scripts\activate # Windows
96
+
97
+ # Install dependencies
98
+ pip install -r requirements_local.txt
99
+ ```
100
+
101
+ ### Step 4: Run the API Server
102
+
103
+ ```bash
104
+ # Start the FastAPI server (runs on port 8001)
105
+ python app_ollama.py
106
+ ```
107
+
108
+ Or using uvicorn directly:
109
+ ```bash
110
+ uvicorn app_ollama:app --host 0.0.0.0 --port 8001
111
+ ```
112
+
113
+ ---
114
+
115
+ ## Testing
116
+
117
+ ### Test 1: Verify Ollama is Running
118
+
119
+ ```bash
120
+ # Check Ollama status
121
+ curl http://127.0.0.1:11434/api/tags
122
+
123
+ # Should return something like:
124
+ # {"models":[{"name":"apple/OpenELM-3B-Instruct"...}]}
125
+ ```
126
+
127
+ ### Test 2: Quick Generation Test
128
+
129
+ ```bash
130
+ # Test basic generation
131
+ curl http://127.0.0.1:11434/api/generate \
132
+ -d '{"model": "apple/OpenELM-3B-Instruct", "prompt": "Say hello!", "stream": false}'
133
+ ```
134
+
135
+ ### Test 3: Run Test Scripts
136
+
137
+ ```bash
138
+ # Make test scripts executable
139
+ chmod +x test_curl.sh
140
+
141
+ # Run curl tests
142
+ ./test_curl.sh
143
+
144
+ # Run Python tests (requires openai package)
145
+ python test_python.py
146
+ ```
147
+
148
+ ### Test 4: Test the Full API Server
149
+
150
+ ```bash
151
+ # Test OpenAI format endpoint
152
+ curl -X POST http://127.0.0.1:8001/v1/chat/completions \
153
+ -H "Content-Type: application/json" \
154
+ -d '{
155
+ "model": "apple/OpenELM-3B-Instruct",
156
+ "messages": [{"role": "user", "content": "Hello!"}],
157
+ "max_tokens": 100
158
+ }'
159
+
160
+ # Test Anthropic format endpoint
161
+ curl -X POST http://127.0.0.1:8001/v1/messages \
162
+ -H "Content-Type: application/json" \
163
+ -d '{
164
+ "model": "apple/OpenELM-3B-Instruct",
165
+ "messages": [{"role": "user", "content": "Hello!"}],
166
+ "max_tokens": 100
167
+ }'
168
+ ```
169
+
170
+ ---
171
+
172
+ ## API Usage
173
+
174
+ ### Using OpenAI SDK (Python)
175
+
176
+ ```python
177
+ from openai import OpenAI
178
+
179
+ # Connect to local Ollama
180
+ client = OpenAI(
181
+ base_url="http://127.0.0.1:11434/v1",
182
+ api_key="ollama", # Any string works
183
+ )
184
+
185
+ # Basic usage
186
+ response = client.chat.completions.create(
187
+ model="apple/OpenELM-3B-Instruct",
188
+ messages=[
189
+ {"role": "system", "content": "You are a helpful assistant."},
190
+ {"role": "user", "content": "Explain quantum computing simply."}
191
+ ],
192
+ max_tokens=200,
193
+ temperature=0.7
194
+ )
195
+
196
+ print(response.choices[0].message.content)
197
+ ```
198
+
199
+ ### Using Anthropic SDK (Python)
200
+
201
+ ```python
202
+ import anthropic
203
+
204
+ # Connect to local Ollama (via API server)
205
+ client = anthropic.Anthropic(
206
+ base_url="http://127.0.0.1:8001/v1",
207
+ api_key="ollama", # Any string works
208
+ )
209
+
210
+ # Basic usage
211
+ message = client.messages.create(
212
+ model="apple/OpenELM-3B-Instruct",
213
+ messages=[{"role": "user", "content": "Hello!"}],
214
+ max_tokens=100
215
+ )
216
+
217
+ print(message.content[0].text)
218
+ ```
219
+
220
+ ### Using cURL
221
+
222
+ ```bash
223
+ # Basic generation
224
+ curl http://127.0.0.1:11434/api/generate \
225
+ -d '{"model": "apple/OpenELM-3B-Instruct", "prompt": "Your prompt here"}'
226
+
227
+ # Chat completion (OpenAI format)
228
+ curl http://127.0.0.1:8001/v1/chat/completions \
229
+ -H "Content-Type: application/json" \
230
+ -d '{
231
+ "model": "apple/OpenELM-3B-Instruct",
232
+ "messages": [{"role": "user", "content": "Your prompt here"}],
233
+ "max_tokens": 100
234
+ }'
235
+ ```
236
+
237
+ ---
238
+
239
+ ## Docker Compose Setup
240
+
241
+ For easier deployment, use Docker Compose:
242
+
243
+ ### Step 1: Start All Services
244
+
245
+ ```bash
246
+ # Start Ollama and API server together
247
+ docker-compose up -d
248
+
249
+ # View logs
250
+ docker-compose logs -f
251
+ ```
252
+
253
+ ### Step 2: Access the Services
254
+
255
+ - **Ollama API**: http://localhost:11434
256
+ - **FastAPI Server**: http://localhost:8001
257
+
258
+ ### Step 3: Stop Services
259
+
260
+ ```bash
261
+ docker-compose down
262
+ ```
263
+
264
+ ---
265
+
266
+ ## Troubleshooting
267
+
268
+ ### Issue: GPU Not Detected
269
+
270
+ **Error**: `Error response from daemon: could not select device driver "" with capabilities: [[gpu]]`
271
+
272
+ **Solution**:
273
+ ```bash
274
+ # Install NVIDIA Container Toolkit
275
+ distribution=$(. /etc/ossa;echo $ID$VERSION_ID)
276
+ curl -s -L https://nvidia.github.io/libnvidia-container/gpgkey | sudo apt-key add -
277
+ curl -s -L https://nvidia.github.io/libnvidia-container/$distribution/libnvidia-container.list | \
278
+ sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
279
+
280
+ sudo apt-get update
281
+ sudo apt-get install -y nvidia-container-toolkit
282
+ sudo systemctl restart docker
283
+ ```
284
+
285
+ ### Issue: Model Download Fails
286
+
287
+ **Error**: `Error: pull model manifest`
288
+
289
+ **Solution**:
290
+ ```bash
291
+ # Check network connection
292
+ curl -I https://huggingface.co
293
+
294
+ # Retry with verbose output
295
+ docker exec -it ollama ollama pull apple/OpenELM-3B-Instruct --verbose
296
+ ```
297
+
298
+ ### Issue: API Server Can't Connect to Ollama
299
+
300
+ **Error**: `Connection refused` or `Ollama not responding`
301
+
302
+ **Solution**:
303
+ ```bash
304
+ # Check if Ollama is running
305
+ docker ps | grep ollama
306
+
307
+ # Check Ollama logs
308
+ docker logs ollama
309
+
310
+ # Restart Ollama
311
+ docker restart ollama
312
+ ```
313
+
314
+ ### Issue: Out of Memory
315
+
316
+ **Error**: `CUDA out of memory`
317
+
318
+ **Solution**:
319
+ - Reduce `max_tokens` parameter
320
+ - Use smaller batch sizes
321
+ - Restart the Ollama container to free memory
322
+
323
+ ### Issue: Port Already in Use
324
+
325
+ **Error**: `Address already in use`
326
+
327
+ **Solution**:
328
+ ```bash
329
+ # Find the process using the port
330
+ lsof -i :11434 # Linux/macOS
331
+ netstat -ano | findstr :11434 # Windows
332
+
333
+ # Kill the process or use a different port
334
+ ```
335
+
336
+ ---
337
+
338
+ ## File Structure
339
+
340
+ ```
341
+ ollama-openelm/
342
+ ├── setup_ollama_openelm.sh # Complete setup script
343
+ ├── app_ollama.py # FastAPI server
344
+ ├── requirements_local.txt # Python dependencies
345
+ ├── docker-compose.yml # Docker Compose configuration
346
+ ├── Dockerfile.api # API server Docker image
347
+ ├── test_python.py # Python test script
348
+ ├── test_curl.sh # cURL test script
349
+ └── README.md # This file
350
+ ```
351
+
352
+ ---
353
+
354
+ ## Environment Variables
355
+
356
+ ### For API Server
357
+
358
+ | Variable | Default | Description |
359
+ |----------|---------|-------------|
360
+ | `OLLAMA_BASE_URL` | `http://127.0.0.1:11434` | Ollama server URL |
361
+ | `OLLAMA_MODEL` | `apple/OpenELM-3B-Instruct` | Model name |
362
+ | `PORT` | `8001` | API server port |
363
+
364
+ ### For Ollama Container
365
+
366
+ | Variable | Description |
367
+ |----------|-------------|
368
+ | `OLLAMA_HOST` | Override the Ollama server URL |
369
+ | `OLLAMA_MODELS` | Path to model storage |
370
+
371
+ ---
372
+
373
+ ## Performance Tips
374
+
375
+ 1. **GPU Memory**: The 3B model uses ~6GB GPU memory
376
+ 2. **CPU Inference**: Falls back to CPU if no GPU available (slower)
377
+ 3. **Batch Size**: Use `num_predict` to control output length
378
+ 4. **Temperature**: Lower values (0.0-0.5) for more deterministic output
379
+
380
+ ---
381
+
382
+ ## Additional Resources
383
+
384
+ - [Ollama Documentation](https://ollama.com/)
385
+ - [OpenELM Model Card](https://huggingface.co/apple/OpenELM-3B-Instruct)
386
+ - [OpenAI API Compatibility](https://platform.openai.com/docs/api-reference)
387
+ - [FastAPI Documentation](https://fastapi.tiangolo.com/)
388
+
389
+ ---
390
+
391
+ ## License
392
+
393
+ This setup is provided for educational and research purposes. The OpenELM models from Apple are released under their respective licenses.
app_ollama.py ADDED
@@ -0,0 +1,391 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """
2
+ OpenELM API Server using Local Ollama
3
+
4
+ This version uses a local Ollama instance instead of Hugging Face,
5
+ providing much faster inference with GPU acceleration.
6
+
7
+ Requirements:
8
+ - Ollama running locally (docker run ollama/ollama)
9
+ - OpenELM model pulled (docker exec ollama ollama pull apple/OpenELM-3B-Instruct)
10
+ - Python packages: pip install -r requirements_local.txt
11
+ """
12
+
13
+ import uuid
14
+ from typing import List, Optional, Dict, Any
15
+ import requests
16
+ from fastapi import FastAPI, HTTPException
17
+ from fastapi.middleware.cors import CORSMiddleware
18
+ from pydantic import BaseModel, Field
19
+ import os
20
+
21
+ # Configuration for local Ollama
22
+ OLLAMA_BASE_URL = os.environ.get("OLLAMA_BASE_URL", "http://127.0.0.1:11434")
23
+ OLLAMA_MODEL = os.environ.get("OLLAMA_MODEL", "apple/OpenELM-3B-Instruct")
24
+
25
+ # Create FastAPI app
26
+ app = FastAPI(
27
+ title="OpenELM API (Ollama)",
28
+ description="OpenAI & Anthropic compatible API using local Ollama instance",
29
+ version="3.0.0"
30
+ )
31
+
32
+ # Add CORS
33
+ app.add_middleware(
34
+ CORSMiddleware,
35
+ allow_origins=["*"],
36
+ allow_credentials=True,
37
+ allow_methods=["*"],
38
+ allow_headers=["*"],
39
+ )
40
+
41
+
42
+ # ==================== Pydantic Models ====================
43
+
44
+ class ChatMessage(BaseModel):
45
+ role: str
46
+ content: str
47
+ name: Optional[str] = None
48
+
49
+
50
+ class ChatCompletionRequest(BaseModel):
51
+ model: str = OLLAMA_MODEL
52
+ messages: List[ChatMessage]
53
+ temperature: Optional[float] = Field(default=None, ge=0.0, le=2.0)
54
+ top_p: Optional[float] = Field(default=None, ge=0.0, le=1.0)
55
+ max_tokens: Optional[int] = Field(default=None, ge=1, le=4096)
56
+ stream: Optional[bool] = False
57
+
58
+
59
+ class ChatCompletionChoice(BaseModel):
60
+ index: int
61
+ message: ChatMessage
62
+ finish_reason: Optional[str] = None
63
+
64
+
65
+ class ChatCompletionUsage(BaseModel):
66
+ prompt_tokens: int
67
+ completion_tokens: int
68
+ total_tokens: int
69
+
70
+
71
+ class ChatCompletionResponse(BaseModel):
72
+ id: str
73
+ object: str = "chat.completion"
74
+ created: int
75
+ model: str
76
+ choices: List[ChatCompletionChoice]
77
+ usage: ChatCompletionUsage
78
+
79
+
80
+ class MessageContent(BaseModel):
81
+ type: str = "text"
82
+ text: str
83
+
84
+
85
+ class Message(BaseModel):
86
+ role: str
87
+ content: str | List[MessageContent]
88
+ name: Optional[str] = None
89
+
90
+
91
+ class Usage(BaseModel):
92
+ input_tokens: int = 0
93
+ output_tokens: int = 0
94
+ total_tokens: int = 0
95
+
96
+
97
+ class ContentBlock(BaseModel):
98
+ type: str = "text"
99
+ text: str
100
+
101
+
102
+ class MessageResponse(BaseModel):
103
+ id: str
104
+ type: str = "message"
105
+ role: str = "assistant"
106
+ content: List[ContentBlock]
107
+ model: str
108
+ stop_reason: Optional[str] = None
109
+ usage: Usage
110
+
111
+
112
+ class MessageCreateParams(BaseModel):
113
+ model: str = OLLAMA_MODEL
114
+ messages: List[Message]
115
+ system: Optional[str] = None
116
+ max_tokens: int = Field(default=1024, ge=1, le=4096)
117
+ temperature: Optional[float] = Field(default=None, ge=0.0, le=1.0)
118
+ stream: Optional[bool] = False
119
+
120
+
121
+ # ==================== Ollama Helper Functions ====================
122
+
123
+ def generate_with_ollama(
124
+ prompt: str,
125
+ system: Optional[str] = None,
126
+ max_tokens: int = 1024,
127
+ temperature: Optional[float] = None,
128
+ stream: bool = False
129
+ ) -> Dict[str, Any]:
130
+ """Generate text using local Ollama instance."""
131
+
132
+ # Build the prompt in chat format
133
+ full_prompt = ""
134
+
135
+ if system:
136
+ full_prompt += f"[System: {system}]\n\n"
137
+
138
+ # Extract messages from prompt
139
+ lines = prompt.split("\n\n")
140
+ for line in lines:
141
+ if line.startswith("User:"):
142
+ full_prompt += f"User: {line[5:].strip()}\n"
143
+ elif line.startswith("Assistant:"):
144
+ full_prompt += f"Assistant: {line[10:].strip()}\n"
145
+ elif line.startswith("User:"):
146
+ full_prompt += f"User: {line[5:].strip()}\n"
147
+
148
+ # Add final assistant prefix
149
+ full_prompt += "Assistant:"
150
+
151
+ # Prepare options
152
+ options = {
153
+ "num_predict": max_tokens,
154
+ }
155
+
156
+ if temperature is not None:
157
+ options["temperature"] = temperature
158
+
159
+ # Make request to Ollama
160
+ response = requests.post(
161
+ f"{OLLAMA_BASE_URL}/api/generate",
162
+ json={
163
+ "model": OLLAMA_MODEL,
164
+ "prompt": full_prompt,
165
+ "stream": stream,
166
+ "options": options
167
+ }
168
+ )
169
+
170
+ if response.status_code != 200:
171
+ raise HTTPException(
172
+ status_code=500,
173
+ detail=f"Ollama request failed: {response.text}"
174
+ )
175
+
176
+ return response.json()
177
+
178
+
179
+ def chat_with_ollama(
180
+ messages: List[ChatMessage],
181
+ max_tokens: int = 1024,
182
+ temperature: Optional[float] = None,
183
+ stream: bool = False
184
+ ) -> Dict[str, Any]:
185
+ """Chat completion using Ollama's chat API."""
186
+
187
+ # Convert messages to Ollama format
188
+ ollama_messages = []
189
+ for msg in messages:
190
+ ollama_messages.append({
191
+ "role": msg.role,
192
+ "content": msg.content
193
+ })
194
+
195
+ # Prepare options
196
+ options = {
197
+ "num_predict": max_tokens,
198
+ }
199
+
200
+ if temperature is not None:
201
+ options["temperature"] = temperature
202
+
203
+ # Make request to Ollama chat API
204
+ response = requests.post(
205
+ f"{OLLAMA_BASE_URL}/v1/chat/completions",
206
+ json={
207
+ "model": OLLAMA_MODEL,
208
+ "messages": ollama_messages,
209
+ "stream": stream,
210
+ "options": options
211
+ }
212
+ )
213
+
214
+ if response.status_code != 200:
215
+ raise HTTPException(
216
+ status_code=500,
217
+ detail=f"Ollama chat request failed: {response.text}"
218
+ )
219
+
220
+ return response.json()
221
+
222
+
223
+ # ==================== API Endpoints ====================
224
+
225
+ @app.get("/", tags=["Root"])
226
+ async def root():
227
+ """Root endpoint with API information."""
228
+ return {
229
+ "name": "OpenELM API (Ollama Local)",
230
+ "version": "3.0.0",
231
+ "model": OLLAMA_MODEL,
232
+ "ollama_url": OLLAMA_BASE_URL,
233
+ "endpoints": {
234
+ "chat": "POST /v1/chat/completions",
235
+ "messages": "POST /v1/messages",
236
+ "health": "GET /health"
237
+ }
238
+ }
239
+
240
+
241
+ @app.get("/health", tags=["Health"])
242
+ async def health_check():
243
+ """Health check endpoint."""
244
+ try:
245
+ response = requests.get(f"{OLLAMA_BASE_URL}/api/tags", timeout=5)
246
+ if response.status_code == 200:
247
+ return {
248
+ "status": "healthy",
249
+ "ollama_connected": True,
250
+ "model": OLLAMA_MODEL
251
+ }
252
+ else:
253
+ return {
254
+ "status": "unhealthy",
255
+ "ollama_connected": False,
256
+ "error": "Ollama not responding"
257
+ }
258
+ except Exception as e:
259
+ return {
260
+ "status": "unhealthy",
261
+ "ollama_connected": False,
262
+ "error": str(e)
263
+ }
264
+
265
+
266
+ @app.post("/v1/chat/completions", response_model=ChatCompletionResponse, tags=["OpenAI"])
267
+ async def create_chat_completion(request: ChatCompletionRequest):
268
+ """Create chat completion (OpenAI API format)."""
269
+
270
+ try:
271
+ # Use Ollama chat API
272
+ result = chat_with_ollama(
273
+ messages=request.messages,
274
+ max_tokens=request.max_tokens or 1024,
275
+ temperature=request.temperature,
276
+ stream=request.stream
277
+ )
278
+
279
+ # Convert to OpenAI format
280
+ choice = result["choices"][0]
281
+ message = choice["message"]
282
+
283
+ response_id = f"chatcmpl-{uuid.uuid4().hex[:12]}"
284
+ timestamp = int(uuid.uuid1().time)
285
+
286
+ return ChatCompletionResponse(
287
+ id=response_id,
288
+ created=timestamp,
289
+ model=OLLAMA_MODEL,
290
+ choices=[
291
+ ChatCompletionChoice(
292
+ index=0,
293
+ message=ChatMessage(role=message["role"], content=message["content"]),
294
+ finish_reason=choice.get("finish_reason", "stop")
295
+ )
296
+ ],
297
+ usage=ChatCompletionUsage(
298
+ prompt_tokens=result["usage"]["prompt_tokens"],
299
+ completion_tokens=result["usage"]["completion_tokens"],
300
+ total_tokens=result["usage"]["total_tokens"]
301
+ )
302
+ )
303
+
304
+ except Exception as e:
305
+ raise HTTPException(status_code=500, detail=f"Generation failed: {str(e)}")
306
+
307
+
308
+ @app.post("/v1/messages", response_model=MessageResponse, tags=["Anthropic"])
309
+ async def create_message(params: MessageCreateParams):
310
+ """Create message (Anthropic API format)."""
311
+
312
+ try:
313
+ # Convert Anthropic messages to prompt
314
+ prompt_parts = []
315
+
316
+ if params.system:
317
+ prompt_parts.append(f"[System: {params.system}]")
318
+
319
+ for msg in params.messages:
320
+ content = msg.content
321
+ if isinstance(content, list):
322
+ content = "".join(b.text for b in content if hasattr(b, 'text'))
323
+
324
+ if msg.role == "user":
325
+ prompt_parts.append(f"User: {content}")
326
+ elif msg.role == "assistant":
327
+ prompt_parts.append(f"Assistant: {content}")
328
+
329
+ prompt_parts.append("Assistant:")
330
+ prompt = "\n\n".join(prompt_parts)
331
+
332
+ # Generate with Ollama
333
+ result = generate_with_ollama(
334
+ prompt=prompt,
335
+ system=params.system,
336
+ max_tokens=params.max_tokens,
337
+ temperature=params.temperature
338
+ )
339
+
340
+ # Extract response
341
+ response_text = result.get("response", "")
342
+
343
+ # Count tokens (approximate)
344
+ input_tokens = len(prompt.split())
345
+ output_tokens = len(response_text.split())
346
+
347
+ return MessageResponse(
348
+ id=f"msg_{uuid.uuid4().hex[:8]}",
349
+ role="assistant",
350
+ content=[ContentBlock(type="text", text=response_text)],
351
+ model=OLLAMA_MODEL,
352
+ stop_reason="end_turn",
353
+ usage=Usage(
354
+ input_tokens=input_tokens,
355
+ output_tokens=output_tokens,
356
+ total_tokens=input_tokens + output_tokens
357
+ )
358
+ )
359
+
360
+ except Exception as e:
361
+ raise HTTPException(status_code=500, detail=f"Generation failed: {str(e)}")
362
+
363
+
364
+ # ==================== Main Entry Point ====================
365
+
366
+ if __name__ == "__main__":
367
+ import uvicorn
368
+
369
+ port = int(os.environ.get("PORT", 8001)) # Different port than Hugging Face Space
370
+
371
+ print(f"""
372
+ ========================================
373
+ OpenELM API Server (Ollama Local)
374
+ ========================================
375
+ Model: {OLLAMA_MODEL}
376
+ Ollama URL: {OLLAMA_BASE_URL}
377
+ Server: http://127.0.0.1:{port}
378
+
379
+ Endpoints:
380
+ OpenAI: POST http://127.0.0.1:{port}/v1/chat/completions
381
+ Anthropic: POST http://127.0.0.1:{port}/v1/messages
382
+ Health: GET http://127.0.0.1:{port}/health
383
+ ========================================
384
+ """)
385
+
386
+ uvicorn.run(
387
+ "app_ollama:app",
388
+ host="0.0.0.0",
389
+ port=port,
390
+ reload=False
391
+ )
docker-compose.yml ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ version: '3.8'
2
+
3
+ services:
4
+ ollama:
5
+ image: ollama/ollama:latest
6
+ container_name: ollama
7
+ ports:
8
+ - "127.0.0.1:11434:11434"
9
+ volumes:
10
+ - ollama_data:/root/.ollama
11
+ deploy:
12
+ resources:
13
+ reservations:
14
+ devices:
15
+ - driver: nvidia
16
+ count: all
17
+ capabilities: [gpu]
18
+ restart: unless-stopped
19
+ healthcheck:
20
+ test: ["CMD", "curl", "-f", "http://localhost:11434/api/tags"]
21
+ interval: 30s
22
+ timeout: 10s
23
+ retries: 3
24
+
25
+ api:
26
+ build:
27
+ context: .
28
+ dockerfile: Dockerfile.api
29
+ container_name: openelm-api
30
+ ports:
31
+ - "8001:8000"
32
+ environment:
33
+ - OLLAMA_BASE_URL=http://ollama:11434
34
+ - OLLAMA_MODEL=apple/OpenELM-3B-Instruct
35
+ depends_on:
36
+ ollama:
37
+ condition: service_healthy
38
+ restart: unless-stopped
39
+
40
+ volumes:
41
+ ollama_data:
requirements_local.txt ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Requirements for Local Ollama Setup
2
+ # Install with: pip install -r requirements_local.txt
3
+
4
+ # FastAPI and server
5
+ fastapi>=0.100.0
6
+ uvicorn[standard]>=0.23.0
7
+
8
+ # HTTP client
9
+ requests>=2.31.0
10
+
11
+ # OpenAI SDK (for testing)
12
+ openai>=1.0.0
13
+
14
+ # For JSON pretty printing (optional)
15
+ pygments>=2.15.0
setup_ollama_openelm.sh ADDED
@@ -0,0 +1,314 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+
3
+ # ============================================
4
+ # Complete Ollama + OpenELM Setup Script
5
+ # ============================================
6
+
7
+ set -e # Exit on error
8
+
9
+ echo "========================================"
10
+ echo "Ollama OpenELM Setup Script"
11
+ echo "========================================"
12
+ echo ""
13
+
14
+ # Color codes for output
15
+ RED='\033[0;31m'
16
+ GREEN='\033[0;32m'
17
+ YELLOW='\033[1;33m'
18
+ NC='\033[0m' # No Color
19
+
20
+ # Function to print colored output
21
+ print_status() {
22
+ echo -e "${GREEN}[✓]${NC} $1"
23
+ }
24
+
25
+ print_warning() {
26
+ echo -e "${YELLOW}[!]${NC} $1"
27
+ }
28
+
29
+ print_error() {
30
+ echo -e "${RED}[✗]${NC} $1"
31
+ }
32
+
33
+ # ============================================
34
+ # Step 1: Check Prerequisites
35
+ # ============================================
36
+ echo "Step 1: Checking prerequisites..."
37
+ echo "-----------------------------------"
38
+
39
+ # Check Docker
40
+ if ! command -v docker &> /dev/null; then
41
+ print_error "Docker is not installed!"
42
+ echo ""
43
+ echo "Please install Docker first:"
44
+ echo " Ubuntu/Debian: sudo apt-get install docker.io"
45
+ echo " macOS: brew install --cask docker"
46
+ echo " Windows: Download from https://docker.com/products/docker-desktop"
47
+ exit 1
48
+ fi
49
+ print_status "Docker is installed"
50
+
51
+ # Check Docker daemon is running
52
+ if ! docker info &> /dev/null; then
53
+ print_error "Docker daemon is not running!"
54
+ echo ""
55
+ echo "Please start Docker:"
56
+ echo " Ubuntu/Debian: sudo systemctl start docker"
57
+ echo " macOS/Windows: Start Docker Desktop application"
58
+ exit 1
59
+ fi
60
+ print_status "Docker daemon is running"
61
+
62
+ # Check NVIDIA driver
63
+ if ! command -v nvidia-smi &> /dev/null; then
64
+ print_warning "NVIDIA driver not found. GPU support may not work!"
65
+ else
66
+ print_status "NVIDIA driver is installed"
67
+ nvidia-smi --query-gpu=name,memory.total --format=csv,noheader 2>/dev/null || echo " GPU detected"
68
+ fi
69
+
70
+ echo ""
71
+
72
+ # ============================================
73
+ # Step 2: Stop and Remove Existing Ollama Container
74
+ # ============================================
75
+ echo "Step 2: Cleaning up existing containers..."
76
+ echo "-----------------------------------"
77
+
78
+ if docker ps -a | grep -q "ollama"; then
79
+ print_warning "Removing existing ollama container..."
80
+ docker stop ollama 2>/dev/null || true
81
+ docker rm ollama 2>/dev/null || true
82
+ print_status "Existing container removed"
83
+ else
84
+ print_status "No existing ollama container found"
85
+ fi
86
+
87
+ echo ""
88
+
89
+ # ============================================
90
+ # Step 3: Start Ollama Container
91
+ # ============================================
92
+ echo "Step 3: Starting Ollama container..."
93
+ echo "-----------------------------------"
94
+
95
+ docker run -d \
96
+ --name ollama \
97
+ -v ollama:/root/.ollama \
98
+ -p 127.0.0.1:11434:11434 \
99
+ --gpus all \
100
+ ollama/ollama
101
+
102
+ print_status "Ollama container started"
103
+ echo "Container ID: $(docker ps -q --filter ancestor=ollama/ollama)"
104
+
105
+ echo ""
106
+
107
+ # ============================================
108
+ # Step 4: Wait for Ollama to be ready
109
+ # ============================================
110
+ echo "Step 4: Waiting for Ollama to initialize..."
111
+ echo "-----------------------------------"
112
+
113
+ max_attempts=30
114
+ attempt=0
115
+ while [ $attempt -lt $max_attempts ]; do
116
+ if docker exec ollama curl -s http://localhost:11434/api/tags > /dev/null 2>&1; then
117
+ print_status "Ollama is ready!"
118
+ break
119
+ fi
120
+
121
+ attempt=$((attempt + 1))
122
+ echo " Attempt $attempt/$max_attempts (waiting 2 seconds)..."
123
+ sleep 2
124
+ done
125
+
126
+ if [ $attempt -eq $max_attempts ]; then
127
+ print_error "Ollama failed to start within 60 seconds"
128
+ echo ""
129
+ echo "Checking logs..."
130
+ docker logs ollama
131
+ exit 1
132
+ fi
133
+
134
+ echo ""
135
+
136
+ # ============================================
137
+ # Step 5: Pull OpenELM Model
138
+ # ============================================
139
+ echo "Step 5: Pulling OpenELM 3B model..."
140
+ echo "-----------------------------------"
141
+ echo "This may take 5-15 minutes depending on your internet connection."
142
+ echo "Model size: approximately 2.1 GB"
143
+ echo ""
144
+
145
+ # Pull the model
146
+ if docker exec -it ollama ollama pull apple/OpenELM-3B-Instruct; then
147
+ print_status "OpenELM model pulled successfully!"
148
+ else
149
+ print_error "Failed to pull OpenELM model"
150
+ exit 1
151
+ fi
152
+
153
+ echo ""
154
+
155
+ # ============================================
156
+ # Step 6: Verify Installation
157
+ # ============================================
158
+ echo "Step 6: Verifying installation..."
159
+ echo "-----------------------------------"
160
+
161
+ echo "Installed models:"
162
+ docker exec ollama ollama list
163
+
164
+ echo ""
165
+ echo "Testing model with a simple prompt..."
166
+ test_response=$(docker exec ollama curl -s http://localhost:11434/api/generate \
167
+ -d '{"model": "apple/OpenELM-3B-Instruct", "prompt": "Hello", "stream": false}')
168
+
169
+ if echo "$test_response" | grep -q "response"; then
170
+ print_status "Model is working!"
171
+ echo ""
172
+ echo "Test response preview:"
173
+ echo "$test_response" | python3 -m json.tool 2>/dev/null | head -20 || echo "$test_response"
174
+ else
175
+ print_warning "Model test returned unexpected response"
176
+ echo "$test_response"
177
+ fi
178
+
179
+ echo ""
180
+
181
+ # ============================================
182
+ # Step 7: Create Test Scripts
183
+ # ============================================
184
+ echo "Step 7: Creating test scripts..."
185
+ echo "-----------------------------------"
186
+
187
+ # Create test_python.py
188
+ cat > test_python.py << 'PYTHON_TEST'
189
+ #!/usr/bin/env python3
190
+ """
191
+ Test script for Ollama OpenELM API
192
+ """
193
+
194
+ from openai import OpenAI
195
+
196
+ def test_ollama():
197
+ """Test Ollama with OpenAI SDK."""
198
+ print("Testing Ollama OpenELM with OpenAI SDK...")
199
+ print("-" * 50)
200
+
201
+ client = OpenAI(
202
+ base_url="http://127.0.0.1:11434/v1",
203
+ api_key="ollama",
204
+ )
205
+
206
+ # Test 1: Basic generation
207
+ print("\n[Test 1] Basic generation:")
208
+ response = client.chat.completions.create(
209
+ model="apple/OpenELM-3B-Instruct",
210
+ messages=[{"role": "user", "content": "Say hello!"}],
211
+ max_tokens=100,
212
+ temperature=0.7
213
+ )
214
+ print(f"Response: {response.choices[0].message.content}")
215
+ print(f"Tokens: {response.usage.total_tokens}")
216
+
217
+ # Test 2: Multi-turn conversation
218
+ print("\n[Test 2] Multi-turn conversation:")
219
+ response = client.chat.completions.create(
220
+ model="apple/OpenELM-3B-Instruct",
221
+ messages=[
222
+ {"role": "user", "content": "What is AI?"},
223
+ {"role": "assistant", "content": "AI is Artificial Intelligence."},
224
+ {"role": "user", "content": "What are examples?"}
225
+ ],
226
+ max_tokens=150,
227
+ temperature=0.5
228
+ )
229
+ print(f"Response: {response.choices[0].message.content}")
230
+
231
+ # Test 3: Creative writing
232
+ print("\n[Test 3] Creative writing:")
233
+ response = client.chat.completions.create(
234
+ model="apple/OpenELM-3B-Instruct",
235
+ messages=[{"role": "user", "content": "Write a short poem about technology."}],
236
+ max_tokens=200,
237
+ temperature=0.8
238
+ )
239
+ print(f"Response: {response.choices[0].message.content}")
240
+
241
+ print("\n" + "=" * 50)
242
+ print("All tests completed successfully!")
243
+
244
+ if __name__ == "__main__":
245
+ test_ollama()
246
+ PYTHON_TEST
247
+
248
+ # Create test_curl.sh
249
+ cat > test_curl.sh << 'CURL_TEST'
250
+ #!/bin/bash
251
+
252
+ # Test curl commands for Ollama OpenELM
253
+
254
+ echo "========================================"
255
+ echo "Testing Ollama OpenELM with curl"
256
+ echo "========================================"
257
+ echo ""
258
+
259
+ # Test 1: Basic generation
260
+ echo "[Test 1] Basic Generation"
261
+ echo "-------------------------"
262
+ curl -s http://127.0.0.1:11434/api/generate \
263
+ -d '{"model": "apple/OpenELM-3B-Instruct", "prompt": "Say hello!", "stream": false}' | \
264
+ python3 -m json.tool 2>/dev/null || echo "Raw output:"
265
+ echo ""
266
+
267
+ # Test 2: With temperature
268
+ echo "[Test 2] Creative Generation (temperature=0.9)"
269
+ echo "------------------------------------------------"
270
+ curl -s http://127.0.0.1:11434/api/generate \
271
+ -d '{"model": "apple/OpenELM-3B-Instruct", "prompt": "Write a short story about AI", "stream": false, "options": {"temperature": 0.9}}' | \
272
+ python3 -m json.tool 2>/dev/null | head -30 || echo "Raw output:"
273
+ echo ""
274
+
275
+ # Test 3: List models
276
+ echo "[Test 3] List Installed Models"
277
+ echo "------------------------------"
278
+ curl -s http://127.0.0.1:11434/api/tags | python3 -m json.tool 2>/dev/null || echo "Raw output:"
279
+ echo ""
280
+
281
+ echo "========================================"
282
+ echo "Tests completed!"
283
+ echo "========================================"
284
+ CURL_TEST
285
+
286
+ chmod +x test_python.py test_curl.sh
287
+ print_status "Test scripts created: test_python.py, test_curl.sh"
288
+
289
+ echo ""
290
+
291
+ # ============================================
292
+ # Summary
293
+ # ============================================
294
+ echo "========================================"
295
+ echo "Setup Complete!"
296
+ echo "========================================"
297
+ echo ""
298
+ echo "Ollama is now running at: http://127.0.0.1:11434"
299
+ echo ""
300
+ echo "Quick Test Commands:"
301
+ echo ""
302
+ echo "# Test with curl:"
303
+ echo "./test_curl.sh"
304
+ echo ""
305
+ echo "# Test with Python OpenAI SDK:"
306
+ echo "pip install openai"
307
+ echo "python test_python.py"
308
+ echo ""
309
+ echo "# Stop Ollama:"
310
+ echo "docker stop ollama"
311
+ echo ""
312
+ echo "# Start Ollama again (model persists):"
313
+ echo "docker start ollama"
314
+ echo ""
test_curl.sh ADDED
@@ -0,0 +1,47 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/bin/bash
2
+
3
+ # Test curl commands for Ollama OpenELM
4
+
5
+ echo "========================================"
6
+ echo "Testing Ollama OpenELM with curl"
7
+ echo "========================================"
8
+ echo ""
9
+
10
+ # Test 1: Basic generation
11
+ echo "[Test 1] Basic Generation"
12
+ echo "-------------------------"
13
+ curl -s http://127.0.0.1:11434/api/generate \
14
+ -d '{"model": "apple/OpenELM-3B-Instruct", "prompt": "Say hello!", "stream": false}' | \
15
+ python3 -m json.tool 2>/dev/null || echo "Raw output:"
16
+ echo ""
17
+
18
+ # Test 2: With temperature
19
+ echo "[Test 2] Creative Generation (temperature=0.9)"
20
+ echo "------------------------------------------------"
21
+ curl -s http://127.0.0.1:11434/api/generate \
22
+ -d '{"model": "apple/OpenELM-3B-Instruct", "prompt": "Write a short story about AI", "stream": false, "options": {"temperature": 0.9}}' | \
23
+ python3 -m json.tool 2>/dev/null | head -30 || echo "Raw output:"
24
+ echo ""
25
+
26
+ # Test 3: List models
27
+ echo "[Test 3] List Installed Models"
28
+ echo "------------------------------"
29
+ curl -s http://127.0.0.1:11434/api/tags | python3 -m json.tool 2>/dev/null || echo "Raw output:"
30
+ echo ""
31
+
32
+ # Test 4: Chat completions format
33
+ echo "[Test 4] Chat Completions Format"
34
+ echo "---------------------------------"
35
+ curl -s http://127.0.0.1:11434/v1/chat/completions \
36
+ -H "Content-Type: application/json" \
37
+ -d '{
38
+ "model": "apple/OpenELM-3B-Instruct",
39
+ "messages": [{"role": "user", "content": "What is 2+2?"}],
40
+ "max_tokens": 50,
41
+ "temperature": 0.0
42
+ }' | python3 -m json.tool 2>/dev/null || echo "Raw output:"
43
+ echo ""
44
+
45
+ echo "========================================"
46
+ echo "Tests completed!"
47
+ echo "========================================"
test_python.py ADDED
@@ -0,0 +1,69 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """
3
+ Test script for Ollama OpenELM API
4
+ """
5
+
6
+ from openai import OpenClient
7
+
8
+ def test_ollama():
9
+ """Test Ollama with OpenAI SDK."""
10
+ print("Testing Ollama OpenELM with OpenAI SDK...")
11
+ print("-" * 50)
12
+
13
+ client = OpenAI(
14
+ base_url="http://127.0.0.1:11434/v1",
15
+ api_key="ollama",
16
+ )
17
+
18
+ # Test 1: Basic generation
19
+ print("\n[Test 1] Basic generation:")
20
+ response = client.chat.completions.create(
21
+ model="apple/OpenELM-3B-Instruct",
22
+ messages=[{"role": "user", "content": "Say hello!"}],
23
+ max_tokens=100,
24
+ temperature=0.7
25
+ )
26
+ print(f"Response: {response.choices[0].message.content}")
27
+ print(f"Tokens: {response.usage.total_tokens}")
28
+
29
+ # Test 2: Multi-turn conversation
30
+ print("\n[Test 2] Multi-turn conversation:")
31
+ response = client.chat.completions.create(
32
+ model="apple/OpenELM-3B-Instruct",
33
+ messages=[
34
+ {"role": "user", "content": "What is AI?"},
35
+ {"role": "assistant", "content": "AI is Artificial Intelligence."},
36
+ {"role": "user", "content": "What are examples?"}
37
+ ],
38
+ max_tokens=150,
39
+ temperature=0.5
40
+ )
41
+ print(f"Response: {response.choices[0].message.content}")
42
+
43
+ # Test 3: Creative writing
44
+ print("\n[Test 3] Creative writing:")
45
+ response = client.chat.completions.create(
46
+ model="apple/OpenELM-3B-Instruct",
47
+ messages=[{"role": "user", "content": "Write a short poem about technology."}],
48
+ max_tokens=200,
49
+ temperature=0.8
50
+ )
51
+ print(f"Response: {response.choices[0].message.content}")
52
+
53
+ # Test 4: Question answering
54
+ print("\n[Test 4] Question answering:")
55
+ response = client.chat.completions.create(
56
+ model="apple/OpenELM-3B-Instruct",
57
+ messages=[{"role": "user", "content": "Explain what is machine learning in simple terms."}],
58
+ max_tokens=250,
59
+ temperature=0.6
60
+ )
61
+ print(f"Response: {response.choices[0].message.content}")
62
+
63
+ print("\n" + "=" * 50)
64
+ print("All tests completed successfully!")
65
+ print("=" * 50)
66
+
67
+
68
+ if __name__ == "__main__":
69
+ test_ollama()