Nile Edge Service
14 min
overview the nile edge service is an add on for existing nile customers, providing operational ownership of the branch wan where a traditional sd wan appliance exposes dozens of operator facing knobs, the nile edge engine collapses that same surface area into a service customers describe the isp circuit they procured, and the service handles path selection, monitoring, steering, ipsec termination, and alerting on their behalf this document describes the capabilities that are available with the service adaptive path selection (aps) application aware, capacity aware steering across multiple isps, with no operator defined thresholds or rules ipsec hub connectivity ikev2 based site to hub tunnels with primary/backup hub definitions and automatic failovers path visibility hop by hop telemetry for the top application in each sla class across the transit isps the service inherits its identity, telemetry, alerting, policy, and support model from the existing nile access service nile access service understands the device/user identity via 802 1x, sso, or fingerprinting passes every session/flow through the built in dpi, which classifies it into an application and the ip sla class it belongs to runs the management plane — policy, telemetry sink, alerting fabric, portal nile edge service uses the data from access service, picks the top app per sla class, and probes/monitors it for latency, packet loss, and jitter computes a mos score from those metrics and factors it into deciding which wan link is best suited for that sla class terminates ipsec tunnels provides content filtering to block urls terminates two or more isp attachments per site (broadband, dia, lte/5g), configured at onboarding onboarding navigate to network setup >> service areas >> edit the site " " >> topology >> " " on the isp icon click on the individual link from the nsb gateway to the isp to add the ip information isp name & circuit number this is used or alerting when the circuits experience an outage or degradation contract start date used to reset the data cap counters mrc to understand the billing and help customers with credits in case of service degradation from the circuit provider link type dropdown style selection to understand if the link is broadband, fiber, dia, lte, satellite, etc contracted link speed used to understand the link utilization of various circuits and feeds into the aps algorithm data cap used to understand the data capacity of the link and feeds into the aps algorithm notably absent from the form sla class definitions, application categorizations, probe endpoints, threshold values, or steering rules the edge service automates the fine tuning of these values dynamically adaptive path selection (aps) aps is the closed control loop that monitors all flows on the branch's wan links, classifies them, measures their quality on every viable path, and steers each class to the best link based on the current conditions the loop runs continuously on the appliance; the customer never tunes a threshold identify sla classification every flow that is passing the nile access service is passed through the dpi engine, this produces an application identity (e g , microsoft teams audio, salesforce api, apple update) ip sla class examples of applications mapped real time zoom, webex, teams, sip, webrtc (voice/video) business critical erp, workday, hr, pos, salesforce transactional web apps, authentication, dashboards, saas (http/2, grpc), ssh, rdp bulk / best effort backups, file sync, software updates, onedrive/dropbox transfers the mapping of sla class to the applications is based on the information extracted in the dpi engine and is managed centrally by the nile cloud service and is based on heuristics of the customer vertical, segment, etc , customers don't need to manually fine tune it for each ip sla class, the edge service picks a representative top application; the application generating the largest share of class volume over a rolling window this is the application used for path monitoring and visibility probe what we measure for each ip sla class, the edge service maintains a per class measurement of three metrics on each isp link latency round trip time (rtt) measured in milliseconds packet loss measured as a percentage over the probe window jitter measured as the inter packet variance the edge service currently does active probing, i e , each edge engine from all the available isp links, generates a probe to the end point identified and deduces the metrics mentioned above the probing cadence today is every 5 mins score mos computation each path/class pair is scored with a mean opinion score (mos) approximation derived from the e model (itu t g 107) with a class specific weighting the resulting mos is on the 1 0 4 5 range a score of 4 0+ is indistinguishable from a good link for the ip sla class; 3 5 4 0 is noticeable degradation but usable for certain sla classes; below 3 5 is where it turns into a user complain qualify dynamic thresholds and link qualification static thresholds do not hold true in todays day and age, a 100ms rtt maybe excellent for a transactional or interactive traffic going to a saas endpoint, but might be unacceptable for the real time traffic the edge service sets its qualification threshold per ip sla class, per site, instead of one number for the whole path when a link drops below threshold for a ip sla class, it's disqualified for that class only; not pulled from service if it still meets the thresholds for other classes, edge keeps steering those applications onto it to avoid flapping between links and resetting active connections, a link isn't marked qualified again until it's held above threshold for 3 consecutive probe cycles capacity awareness quality (mos) is only one half of the steering decision the other half is capacity, and capacity awareness begins the moment a circuit is onboarded circuit level inputs at onboarding for every isp circuit, the admin enters the contracted upload speed, contracted download speed, and (if applicable) the data cap, as part of the single isp onboarding form continuous utilization tracking once the circuit is live, the edge service continuously measures how much of that contracted upload and download capacity is actually being consumed, in both directions, in real time volume by ip sla class because every flow is already classified into an sla class (real time, interactive, bulk, background) by the shared dpi engine, aps also knows which classes are driving that utilization in practice, background/bulk traffic (backups, updates, file sync) tends to be the heavy consumer of both bandwidth and data cap, while real time and interactive traffic (voice, video, saas, transactional apps) typically consumes comparatively little feeding this into link selection this combination — contracted capacity from onboarding, current utilization, and per class volume, becomes one of the inputs to the steering algorithm, alongside link quality (mos) before steering an ip sla class to a given link, we check whether that link has enough headroom against its contracted speed (and, where a data cap exists, against remaining cap) to take that class's traffic if not, the class is steered to the alternative link instead data cap protection because bulk class traffic is the class most likely to erode a data cap, this is where the capacity input has the most visible effect as cumulative usage approaches the cap, bulk traffic is the first to be moved off that circuit the net effect capacity is not a separate rule engine bolted onto quality scoring; it's a per circuit, per direction, per class decision built from data the customer already gave nile at onboarding, kept current in real time, and used as one of the deciding factors (alongside mos) for which link a given sla class should ride steering the control loop produces one decision per (ip sla class, path) pair every 5 mins the decision is applied to new flows for that class existing flows are sticky to their original path until either the path is fully disqualified for the class (hard re steer; existing flows are steered, disrupting the connections) if the link is down, i e , port or hardware failure failure modes and fallback behavior scenario behaviour nsb cannot reach nile cloud steering continues on last known policy; telemetry buffers and replays on reconnect all paths degraded for a class aps picks the least bad path (highest current mos) and raises an alert probing inconclusive for ≥ 2 cycles the (class, path) pair is treated as qualified — steering does not preemptively disqualify on absence of evidence single isp at a site aps still runs the loop; the path selection output is trivial, but classification, capacity awareness and alerting all remain path visibility for the top application in each ip sla class, the nile edge service probes every hop on the path from the branch to the application's endpoint the result is per hop telemetry — ip, bgp as, rtt, loss, jitter — usable for triage, observability and customer facing diagnostics