Connect your software
Keep your model serving and run the first test on the same machine as Link.
Send a local request
Slide 1 of 4: Model Serving in Control.
Read the API steps and copy commands
Send a request to the model running through Link, then use the same endpoint in your application.
Before you start
Complete setup and serving. Keep Link running and open Terminal on the same Mac. Aventail setup serves on port 8100. If you configured another endpoint, use its actual port.
Tested: Aventail Link 2.0.0 · Llama 3.2 3B Instruct Q4_K_M · Apple Silicon · 2 October 2026.
1. Find your model ID
Replace 8100 if you chose a different port:
curl --fail-with-body --silent --show-error \
http://127.0.0.1:8100/v1/models
Copy an id from the returned data array. Use that exact value, including spaces and punctuation, in the next request.
Use 127.0.0.1 for requests on this Mac, not the bind address 0.0.0.0. Serving can listen on all interfaces; review network exposure before connecting other machines.
2. Send a request
Save this as request.json, replacing the model ID with yours:
{
"model": "Llama-3.2-3B-Instruct-Q4_K_M",
"messages": [{"role": "user", "content": "Reply with exactly: Aventail is ready."}],
"max_tokens": 32,
"temperature": 0,
"stream": false
}
Run:
curl --fail-with-body --silent --show-error --max-time 180 \
-H 'Content-Type: application/json' \
--data-binary @request.json \
http://127.0.0.1:8100/v1/chat/completions
3. Check the response
A successful test needs a successful HTTP response and non-empty choices[0].message.content. The recorded response included:
{
"choices": [
{
"finish_reason": "stop",
"index": 0,
"message": {
"role": "assistant",
"content": "Aventail is ready."
}
}
]
}
The wording can vary. The response may report the model file path instead of the request's alias.
Connect an application
Use http://127.0.0.1:8100/v1 as the base URL for an OpenAI-compatible client on this Mac, with the model ID returned above. Browser clients also need an allowed origin; see the Link integration reference.
For app-specific setup, see Continue in VS Code, Open WebUI, Obsidian, or OpenClaw. You can also chat in Workspace.
If the request fails
| Symptom | Next action |
|---|---|
| Connection refused | Confirm Link is running, serving started, and the port matches Control. |
| Model-name/router error | Call /v1/models again and copy an exact id. |
| HTTP success but empty content | Inspect finish_reason. If it is length, increase max_tokens for the model and task. |
| First request is slow | Allow time for engine/model startup. Retain any error text rather than repeatedly redeploying the model. |
To stop serving, open Serving in Control, expand the device and turn off the model's switch. The model stays installed.