# Routing

Source: [https://docs.aventail.co.uk/docs/fleet/routing](https://docs.aventail.co.uk/docs/fleet/routing)

Use **Fleet operations → Routing** to choose where devices run inference. Follow the screenshots to give one test device an **on-device route with a cloud fallback**, then use the written guide for field explanations and verification.

### Route locally, with a cloud fallback

Development Control, 2 October 2026. Keys and identifying details are obscured.

1.  **Open Routing**
    
    Select the intended policy. The page brings together policy rules, provider connections and each device’s route and delivery status.
    
    [Screenshot: Policy rules and default route — close-up from Fleet Routing page.](https://docs.aventail.co.uk/img/guides/context/fleet-routing.png)
    
2.  **Prepare the cloud connection**
    
    Reuse an approved connection in API key vault, or choose Add key and provide its name, provider and API key. Select a provider-supported model.
    
    This form was inspected without adding another key; the walkthrough used an existing connection.
    
    [Screenshot: Add API key form with Google Gemini selected and the key field empty.](https://docs.aventail.co.uk/img/control/fleet-2026-10-02/routing-provider-form.png)
    
3.  **Tag the test device**
    
    Open the device and add a dedicated tag. Keep policy assignment on Auto for this example. Check that no other devices share the test tag.
    
    [Screenshot: Device tags and automatic policy assignment for Documentation fallback.](https://docs.aventail.co.uk/img/control/fleet-2026-10-02/routing-device-tag.png)
    
4.  **Create a policy for that tag**
    
    Choose New policy, enter a name and select the device’s tag. Leave Inherit Org default rules off to keep this example isolated.
    
    [Screenshot: New routing policy dialog with Documentation fallback and docs-walkthrough selected.](https://docs.aventail.co.uk/img/control/fleet-2026-10-02/routing-new-policy.png)
    
5.  **Add the local-first rule**
    
    Choose a condition the test device meets, an On device destination and a cloud fallback hop. Check the model and matching-device count before saving.
    
    CPU > 0 is a test condition. The written reference explains other signals, hold periods and recovery settings.
    
    [Screenshot: Complete new routing rule form: CPU above zero, instant hold, On device primary and Google Gemini fallback.](https://docs.aventail.co.uk/img/control/fleet-2026-10-02/routing-local-cloud-rule.png)
    
6.  **Check the saved rule and default**
    
    The first matching enabled rule wins. This rule tries On device, then its Gemini fallback on failure. The ELSE route remains On device.
    
    [Screenshot: Saved local-first routing rule, fallback destination, enabled switch and On device default route.](https://docs.aventail.co.uk/img/control/fleet-2026-10-02/routing-rule-order.png)
    
7.  **Wait for the device to confirm**
    
    Check Route, Delivery and Signal. This device reports applied and names the local-first rule. A policy acknowledgement still needs an inference test.
    
    [Screenshot: Documentation Mac row showing online, the matching routing rule, applied delivery and CPU signal.](https://docs.aventail.co.uk/img/control/fleet-2026-10-02/routing-delivery-applied.png)
    
8.  **Verify both inference paths**
    
    Send a harmless request locally, then test the fallback with local serving briefly stopped. Restore serving afterward. Check the responses and the destinations in Recent usage.
    
    Both requests succeeded in this walkthrough. The written test includes the request and observed cloud model.
    
    [Screenshot: Device inference routing panel showing On device as effective route and both On device and Google Gemini usage.](https://docs.aventail.co.uk/img/control/fleet-2026-10-02/routing-recent-usage.png)
    

Cloud requests leave the device

A local API address does not guarantee local processing when a cloud route or fallback is configured. Requests sent to the provider leave the device and may incur charges. Use an approved connection and suitable test data.

**Read the routing setup steps**

## Before you start

You need **Fleet Admin** or **Super Admin** access to manage policies and provider keys. Prepare an online, enrolled test device with Aventail Link and a local model that already answers requests. See [Add devices](https://docs.aventail.co.uk/docs/fleet/overview) and [Connect your software](https://docs.aventail.co.uk/docs/getting-started/run-first-interface).

The example uses a device named **Documentation Mac**, the tag `docs-walkthrough`, a policy named **Documentation fallback**, and an existing Google Gemini connection. Use your own names and a model available to your provider account. Start with one device so the policy's scope is easy to check.

## Prepare a cloud connection

Open **API key vault** on the Routing page. An existing approved connection can be reused; its **Test** control checks the provider key. A successful key test is not a test of your routing rule.

To add a connection:

1.  Select **Add key** and give it a recognisable **Key name**.
2.  Choose the **Provider**.
3.  Optionally select a **Default model**, or enter a model ID supported by that provider.
4.  Enter the provider's API key and select **Add to vault**. Check the result before using the connection in a rule.

A provider API key authorises cloud requests. It is different from an [enrollment key](https://docs.aventail.co.uk/docs/fleet/keys-and-analytics), which lets devices join the fleet. Keep the full value out of screenshots and shared commands. The screenshot shows the form with the key field empty; the tested walkthrough reused an existing connection.

## Choose the devices

1.  Open the test device from **Devices** or the Routing table. Under **Tags**, select **new tag** to create a descriptive tag, or select an existing tag.
2.  Return to **Routing → New policy**. Enter the policy name and select the tag under **Auto-assign tags**.
3.  Leave **Inherit Org default rules** off for this isolated example, then select **Create policy**.
4.  Select the new policy and check its device count. Open the test device and confirm the **Policy** section shows the expected policy and tag assignment.
5.  Keep **Pin to a policy** set to **Auto (follow tag selectors)** and **Manual route** set to **Follow policy (no manual route)** for this test.

Selecting a tag is required in the policy-creation flow shown here. Changes to device tags can immediately change which policy governs them. Do not reuse a tag that would unintentionally include other devices.

## Add a local-first rule

In the new policy, leave **Default route** as **On device**, then select **New rule**.

| Field | Value used in this walkthrough | Purpose |
| --- | --- | --- |
| Rule name | Local first, cloud on failure | Makes the rule easy to identify in device status |
| Condition | CPU (%) **\> 0** | Matches the active test Mac; this is not an unconditional rule |
| Conditions must hold for | **0** minutes | Match immediately when the condition is met |
| Route matched devices to | **On device** | Let the agent choose the local model |
| When a device stops matching | **Leaves the group** | Re-evaluate the rules when the condition clears |
| Fallback hop 1 | Approved Google Gemini connection | Try the provider if the local destination fails |
| Fallback model | A model listed for that key | The tested connection used `gemini-3.6-flash` |

Select **fallback hop** to add the fallback destination, then choose its model. Check the sentence showing how many devices the rule would match and select **Save rule**.

The saved rule should appear above the **ELSE** default route, with its switch enabled and its cloud destination shown after **on failure**. Confirm the policy still governs only the intended devices.

**Read how policies, conditions and fallbacks work**

## Which policy governs a device?

Policy selection and rule evaluation are separate:

1.  An explicit **Pin to a policy** takes priority over tag assignment.
2.  Otherwise, matching tags select a policy. If retagging creates an overlap, policy evaluation order resolves it; use the policy move controls to review precedence.
3.  Devices with no matching policy use **Org default**.

Within the selected policy, enabled rules are evaluated from top to bottom. The first match wins. If **Inherit Org default rules** is on, the inherited rules are evaluated after the policy's own rules. The **ELSE** default route catches devices for which no rule matches.

A device's **Manual route** bypasses this rule evaluation. Return it to **Follow policy (no manual route)** when the override is no longer needed.

## Conditions and the hold period

Use **AND condition** when every condition in a group must be true. Use **OR group** when any one group may trigger the rule. For example, high CPU **and** low available RAM require both conditions in one group; high temperature **or** low disk space use separate groups.

Available signals include CPU, GPU, RAM headroom, temperature, disk space, power source, LLM speed and time to first token. Use the units shown beside the field. Model metrics refer to the model currently serving that modality; an idle device may have no recent inference measurements.

The hold period accepts whole minutes from **0 to 1440**. Zero is instant. A longer hold avoids switching routes for a brief spike; it runs on the device and is limited by the metrics sampling interval.

## Recovery and sticky routes

| Exit setting | When the condition clears |
| --- | --- |
| Leaves the group | The device returns to rule evaluation immediately |
| Sticky | The destination remains until the rule is edited or removed; a higher-priority matching rule can still win |

Saving any rule change—including a name change—restarts its hold clock and releases its sticky state on devices. Review the effect before editing a rule during an active workload.

## Destinations and fallback order

**On device** leaves model choice to the agent. A **provider connection** uses the cloud model selected in the rule. A **network peer** must offer the selected model and be available to serve other devices; the device drawer has a separate **Serve other devices** control.

A fallback chain has up to **three hops**, tried in order after a request error or timeout, authentication failure, provider outage, or unavailable destination. Network peers are not fallback hops. Slow responses alone do not trigger fallback; use a suitable performance condition if you want slowness to select a different primary route.

If every destination fails, the chain is retried from the beginning with backoff. The interface specifies two passes for interactive requests; background jobs continue cycling until a destination recovers. A fallback is a recovery path, not a guarantee that a request will succeed.

**Read the routing tests and troubleshooting guide**

## Check delivery and requests

The Routing table separates three useful checks:

| Column | What it tells you |
| --- | --- |
| Route | The effective destination and the policy, rule or default responsible |
| Delivery | Whether the device confirmed the routing policy |
| Signal | The reported condition behind the decision, or why rule state is unknown |

Wait for **applied** before testing. **Delivering** means the policy is still on its way. **Unconfirmed** means the config update completed but this policy revision was not confirmed; inspect the agent version and connection before relying on it. An offline device's old route is not evidence that it can process requests now.

### Verify local inference

Use the model ID returned by the device's `/v1/models` endpoint. This example uses the walkthrough's `Gemma4`:

```bash
curl --fail-with-body http://127.0.0.1:8100/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Gemma4",
    "messages": [{"role": "user", "content": "Reply with exactly: Fleet enrollment verified."}],
    "max_tokens": 128,
    "temperature": 0
  }'
```

With local serving available, check for a completed assistant answer and inspect **Recent usage** in the device drawer for on-device activity. **Applied**, **online** and **serving** are useful state checks; none replaces a successful response.

### Verify the cloud fallback

On the isolated test device, briefly stop serving the local model, then repeat the same harmless request. Check the response's model and the device's **Recent usage** for the provider destination. Restore local serving and repeat the request to confirm local operation again.

In the recorded test, the local request succeeded; with Gemma4 stopped, the same request returned **HTTP 200**, model `gemini-3.6-flash`, and the answer `Fleet enrollment verified.` Local serving was restored afterward. The screenshots show both destinations in usage; use the response and destination together, rather than treating an aggregate counter as an exact request count.

## Troubleshooting

| Symptom | Check next |
| --- | --- |
| Create policy is unavailable | Enter a name and select an auto-assignment tag. Check whether another policy already claims devices with that tag. |
| A different policy is active | Inspect the device's pin, tags and policy precedence. |
| The rule does not act | Check its enabled switch, order, current signal, comparison, units and hold period. An earlier rule may win. |
| Editing the policy changes nothing on a device | Check for a manual route and incomplete policy delivery. |
| The destination remains after recovery | Check **Sticky** exit behaviour. Review the impact before editing or removing the rule. |
| Cloud fallback is not used for a slow response | Slowness alone is not a fallback trigger. Check whether a performance condition should change the primary route instead. |
| The provider rejects the request | Test the vault connection; check the selected model, provider access and quota. |
| A peer reports **Model not offered** | Check that the peer offers the exact model and allows peer serving. |
| Every destination fails | Inspect errors at each destination, restore one working route and retry a harmless request. |

Next: [manage team access](https://docs.aventail.co.uk/docs/fleet/team) or [monitor devices](https://docs.aventail.co.uk/docs/fleet/dashboard).
