Skip to content

Transport and limits ​

xgress3 relays HTTP over a single outbound tunnel from your agent to the regional gateway. Multiple client requests share that tunnel at the same time. Gateway and agent must be deployed together at the same version — mixed versions are not supported in production.

This page describes service ingress rate limits, what you can configure on the agent, what behavior to expect, and errors you may see.

For ingress authentication and hostnames, see Overview and Ingress authentication.


How traffic flows ​

  1. The client calls your public service hostname (https://{account}--{service}.{region}.xg3.io/...).
  2. The gateway authenticates the request and checks service policy.
  3. The agent receives the relayed request on the tunnel and forwards it to your backend.
  4. The response returns on the same path. Every response includes X-Xg3-Request-Id.

Agent limits (operator-controlled) ​

These apply on your agent host and are the main throughput knob you control.

LimitDefaultConfigure
Concurrent backend requests128 per tunnelXG3_AGENT_MAX_CONCURRENT_BACKEND_REQUESTS
bash
export XG3_AGENT_MAX_CONCURRENT_BACKEND_REQUESTS=256

Raise this only when your backends can absorb additional parallel connections and CPU. Invalid or non-positive values fall back to 128.

When the agent cannot accept another concurrent forward, or a backend forward fails unexpectedly, the gateway returns 502 with Content-Type: application/vnd.xg3.error+json and a message such as Concurrent backend request limit exceeded.

See Agent operations for tunnel behavior, reconnection, and exit codes.


Service ingress rate limits ​

Tenant traffic to https://{account}--{service}.{region}.xg3.io/... is rate-limited per service hostname. Limits are not per client IP or per credential: every caller of that hostname shares one bucket.

Default
Burst1000 requests
Sustained refill300 requests per minute

After a burst, new requests are admitted as tokens refill (on average five per second). A client that stays at or below 300 requests/minute after the burst will not be limited. A client that keeps sending faster than the refill rate will receive 429 until it slows down.

When a request is limited ​

The gateway returns:

Value
HTTP status429
X-Xg3-StatusRateLimited
Retry-AfterWhole seconds until another single request may be allowed (not the time to refill the full burst). At the default refill rate this is typically 1.
X-Xg3-Request-IdCorrelation id (also in the JSON body)
Content-Typeapplication/vnd.xg3.error+json
json
{
  "error": {
    "code": "XG3_RATE_LIMIT_EXCEEDED",
    "message": "Rate limit exceeded",
    "request_id": "01J…"
  }
}

Honor Retry-After before sending another request. Do not busy-loop. Retry-After is delay-seconds (RFC 7231), minimum 1. It does not mean the burst of 1000 has been restored — only that one more request may succeed.

How to build against this ​

  1. Prefer staying under 300 requests/minute per service hostname in steady state.
  2. Short bursts up to 1000 are allowed; after that, wait for refill.
  3. On 429, wait Retry-After seconds, then retry that one request. If you still receive 429, wait again using the new header.
  4. Spread load across different services if you need higher aggregate throughput — each hostname has its own bucket.

Enrollment, certificate, and OAuth calls on the regional apex (https://{region}.xg3.io) use separate limits — see Platform API rate limits.


Parallel downloads and response pacing ​

All relayed requests share one tunnel. xgress3 paces each response independently: a slow or large download should not block other requests on the same service from starting or receiving their response headers.

What you should observe

  • Several downloads or API calls in parallel complete without one large file stalling small requests for many seconds in the browser (“Waiting for server response”).
  • Cancelling or closing a large download mid-stream should not permanently reduce how many concurrent requests the agent can handle.

If you still see long delays on small requests while a large download runs

  1. Confirm the agent is on a current release.
  2. Check the agent is connected in the console and the service is enabled.
  3. Review agent logs under {data-dir}/logs/ — the agent may have received requests while the client still waited at the gateway.
  4. Contact support with X-Xg3-Request-Id from a slow request.

Ingress errors you may see ​

These are customer-visible outcomes.

HTTPX-Xg3-Status or bodyTypical cause
429RateLimitedToo many requests to this service — see Service ingress rate limits; honor Retry-After
502Policy / validation codesAuth deny, headers too large, relay error — see Deny codes
502Agent error JSONAgent concurrency exceeded or backend forward failed
504AGENT_UNAVAILABLENo active tunnel (agent offline or reconnecting)

The gateway validates caller identity and service policy before relaying. The agent does not validate JWT or service keys.


Platform API rate limits ​

Enrollment and token endpoints on the regional apex (https://{region}.xg3.io) are rate-limited separately from service ingress. If you automate enrollment or token minting, honor Retry-After on 429 (typically 60 seconds on these endpoints) and avoid tight polling loops.

EndpointGuidance
POST /enrollOne enrollment flow per agent; do not burst
GET /enroll/statusPoll at reasonable intervals (agent handles enrollment internally in normal installs)
POST /oauth/tokenCache tokens until near expiry; do not mint on every API call

Production agents use the tunnel for renewal — not repeated CSR REST calls.


Request logging limits ​

When Request logging is enabled on a service, Request Lookup retains metadata for 7 days.

When detailed request logging is included in your plan and you enable body capture, xgress3 stores request and backend response bodies up to 1 MB each per request; larger bodies are truncated. Request-body capture applies only when the request is relayed to your backend.

See Request logging.


Tunnel keepalive and reconnection ​

  • The agent sends application heartbeats on the tunnel so the connection stays active during idle periods.
  • The gateway and network stack use transport keepalive to detect abrupt disconnects (for example unplugged cable) within tens of seconds.
  • On tunnel loss, the agent reconnects with exponential backoff (1 s → 120 s cap, resetting to 1 s after a prior successful connection). No operator action is required.

Details: Agent operations — Reconnection. Planned regional updates use the same reconnect path — Cloud and regions — Planned maintenance.


What is not limited on the agent side ​

  • Backend HTTP timeout: the agent does not impose a fixed timeout on calls to your backend; long-running backend work continues until the client or gateway ends the relay.
  • Request body size: xgress3 streams request bodies end-to-end and does not impose a platform body-size ceiling. Large uploads are limited only by your backend (and any proxy in front of it), not by a fixed megabyte cap at the tenant hostname.

Ensure your backends and any upstream proxies tolerate the concurrency you configure on the agent.