Skip to main content

Routing

Use Fleet operations → Routing to choose where devices run inference. Follow the screenshots to give one test device an on-device route with a cloud fallback, then use the written guide for field explanations and verification.

AVENTAIL / FLEET

Route locally, with a cloud fallback

Full view

Fleet Routing page

Routing brings together policies, provider connections and device delivery status.

Slide 1 of 10: Fleet Routing page.

Development Control, 2 October 2026. Keys and identifying details are obscured. Use the arrows or progress marks to move between slides.
AVENTAIL / FLEET

Route locally, with a cloud fallback

Full view

Fleet Routing page

Routing brings together policies, provider connections and device delivery status.

Slide 1 of 10: Fleet Routing page.

Cloud requests leave the device

A local API address does not guarantee local processing when a cloud route or fallback is configured. Requests sent to the provider leave the device and may incur charges. Use an approved connection and suitable test data.

Read the routing setup steps

Before you start​

You need Fleet Admin or Super Admin access to manage policies and provider keys. Prepare an online, enrolled test device with Aventail Link and a local model that already answers requests. See Add devices and Connect your software.

The example uses a device named Documentation Mac, the tag docs-walkthrough, a policy named Documentation fallback, and an existing Google Gemini connection. Use your own names and a model available to your provider account. Start with one device so the policy's scope is easy to check.

Prepare a cloud connection​

Open API key vault on the Routing page. An existing approved connection can be reused; its Test control checks the provider key. A successful key test is not a test of your routing rule.

To add a connection:

  1. Select Add key and give it a recognisable Key name.
  2. Choose the Provider.
  3. Optionally select a Default model, or enter a model ID supported by that provider.
  4. Enter the provider's API key and select Add to vault. Check the result before using the connection in a rule.

A provider API key authorises cloud requests. It is different from an enrollment key, which lets devices join the fleet. Keep the full value out of screenshots and shared commands. The screenshot shows the form with the key field empty; the tested walkthrough reused an existing connection.

Choose the devices​

  1. Open the test device from Devices or the Routing table. Under Tags, select new tag to create a descriptive tag, or select an existing tag.
  2. Return to Routing → New policy. Enter the policy name and select the tag under Auto-assign tags.
  3. Leave Inherit Org default rules off for this isolated example, then select Create policy.
  4. Select the new policy and check its device count. Open the test device and confirm the Policy section shows the expected policy and tag assignment.
  5. Keep Pin to a policy set to Auto (follow tag selectors) and Manual route set to Follow policy (no manual route) for this test.

Selecting a tag is required in the policy-creation flow shown here. Changes to device tags can immediately change which policy governs them. Do not reuse a tag that would unintentionally include other devices.

Add a local-first rule​

In the new policy, leave Default route as On device, then select New rule.

FieldValue used in this walkthroughPurpose
Rule nameLocal first, cloud on failureMakes the rule easy to identify in device status
ConditionCPU (%) > 0Matches the active test Mac; this is not an unconditional rule
Conditions must hold for0 minutesMatch immediately when the condition is met
Route matched devices toOn deviceLet the agent choose the local model
When a device stops matchingLeaves the groupRe-evaluate the rules when the condition clears
Fallback hop 1Approved Google Gemini connectionTry the provider if the local destination fails
Fallback modelA model listed for that keyThe tested connection used gemini-3.6-flash

Select fallback hop to add the fallback destination, then choose its model. Check the sentence showing how many devices the rule would match and select Save rule.

The saved rule should appear above the ELSE default route, with its switch enabled and its cloud destination shown after on failure. Confirm the policy still governs only the intended devices.

Read how policies, conditions and fallbacks work

Which policy governs a device?​

Policy selection and rule evaluation are separate:

  1. An explicit Pin to a policy takes priority over tag assignment.
  2. Otherwise, matching tags select a policy. If retagging creates an overlap, policy evaluation order resolves it; use the policy move controls to review precedence.
  3. Devices with no matching policy use Org default.

Within the selected policy, enabled rules are evaluated from top to bottom. The first match wins. If Inherit Org default rules is on, the inherited rules are evaluated after the policy's own rules. The ELSE default route catches devices for which no rule matches.

A device's Manual route bypasses this rule evaluation. Return it to Follow policy (no manual route) when the override is no longer needed.

Conditions and the hold period​

Use AND condition when every condition in a group must be true. Use OR group when any one group may trigger the rule. For example, high CPU and low available RAM require both conditions in one group; high temperature or low disk space use separate groups.

Available signals include CPU, GPU, RAM headroom, temperature, disk space, power source, LLM speed and time to first token. Use the units shown beside the field. Model metrics refer to the model currently serving that modality; an idle device may have no recent inference measurements.

The hold period accepts whole minutes from 0 to 1440. Zero is instant. A longer hold avoids switching routes for a brief spike; it runs on the device and is limited by the metrics sampling interval.

Recovery and sticky routes​

Exit settingWhen the condition clears
Leaves the groupThe device returns to rule evaluation immediately
StickyThe destination remains until the rule is edited or removed; a higher-priority matching rule can still win

Saving any rule change—including a name change—restarts its hold clock and releases its sticky state on devices. Review the effect before editing a rule during an active workload.

Destinations and fallback order​

On device leaves model choice to the agent. A provider connection uses the cloud model selected in the rule. A network peer must offer the selected model and be available to serve other devices; the device drawer has a separate Serve other devices control.

A fallback chain has up to three hops, tried in order after a request error or timeout, authentication failure, provider outage, or unavailable destination. Network peers are not fallback hops. Slow responses alone do not trigger fallback; use a suitable performance condition if you want slowness to select a different primary route.

If every destination fails, the chain is retried from the beginning with backoff. The interface specifies two passes for interactive requests; background jobs continue cycling until a destination recovers. A fallback is a recovery path, not a guarantee that a request will succeed.

Read the routing tests and troubleshooting guide

Check delivery and requests​

The Routing table separates three useful checks:

ColumnWhat it tells you
RouteThe effective destination and the policy, rule or default responsible
DeliveryWhether the device confirmed the routing policy
SignalThe reported condition behind the decision, or why rule state is unknown

Wait for applied before testing. Delivering means the policy is still on its way. Unconfirmed means the config update completed but this policy revision was not confirmed; inspect the agent version and connection before relying on it. An offline device's old route is not evidence that it can process requests now.

Verify local inference​

Use the model ID returned by the device's /v1/models endpoint. This example uses the walkthrough's Gemma4:

curl --fail-with-body http://127.0.0.1:8100/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "Gemma4",
"messages": [{"role": "user", "content": "Reply with exactly: Fleet enrollment verified."}],
"max_tokens": 128,
"temperature": 0
}'

With local serving available, check for a completed assistant answer and inspect Recent usage in the device drawer for on-device activity. Applied, online and serving are useful state checks; none replaces a successful response.

Verify the cloud fallback​

On the isolated test device, briefly stop serving the local model, then repeat the same harmless request. Check the response's model and the device's Recent usage for the provider destination. Restore local serving and repeat the request to confirm local operation again.

In the recorded test, the local request succeeded; with Gemma4 stopped, the same request returned HTTP 200, model gemini-3.6-flash, and the answer Fleet enrollment verified. Local serving was restored afterward. The screenshots show both destinations in usage; use the response and destination together, rather than treating an aggregate counter as an exact request count.

Troubleshooting​

SymptomCheck next
Create policy is unavailableEnter a name and select an auto-assignment tag. Check whether another policy already claims devices with that tag.
A different policy is activeInspect the device's pin, tags and policy precedence.
The rule does not actCheck its enabled switch, order, current signal, comparison, units and hold period. An earlier rule may win.
Editing the policy changes nothing on a deviceCheck for a manual route and incomplete policy delivery.
The destination remains after recoveryCheck Sticky exit behaviour. Review the impact before editing or removing the rule.
Cloud fallback is not used for a slow responseSlowness alone is not a fallback trigger. Check whether a performance condition should change the primary route instead.
The provider rejects the requestTest the vault connection; check the selected model, provider access and quota.
A peer reports Model not offeredCheck that the peer offers the exact model and allows peer serving.
Every destination failsInspect errors at each destination, restore one working route and retry a harmless request.

Next: manage team access or monitor devices.