Appearance
Transport and limits
xgress3 relays HTTP over a single outbound tunnel from your agent to the regional gateway. Multiple client requests share that tunnel at the same time. Gateway and agent must be deployed together at the same version — mixed versions are not supported in production.
This page describes service ingress rate limits, what you can configure on the agent, what behavior to expect, and errors you may see.
For ingress authentication and hostnames, see Overview and Ingress authentication.
How traffic flows
- The client calls your public service hostname (
https://{account}--{service}.{region}.xg3.io/...). - The gateway authenticates the request and checks service policy.
- The agent receives the relayed request on the tunnel and forwards it to your backend.
- The response returns on the same path. Every response includes
X-Xg3-Request-Id.
Agent limits (operator-controlled)
These apply on your agent host and are the main throughput knob you control.
| Limit | Default | Configure |
|---|---|---|
| Concurrent backend requests | 128 per tunnel | XG3_AGENT_MAX_CONCURRENT_BACKEND_REQUESTS |
bash
export XG3_AGENT_MAX_CONCURRENT_BACKEND_REQUESTS=256Raise this only when your backends can absorb additional parallel connections and CPU. Invalid or non-positive values fall back to 128.
When the agent cannot accept another concurrent forward, or a backend forward fails unexpectedly, the gateway returns 502 with Content-Type: application/vnd.xg3.error+json and a message such as Concurrent backend request limit exceeded.
See Agent operations for tunnel behavior, reconnection, and exit codes.
Service ingress rate limits
Tenant traffic to https://{account}--{service}.{region}.xg3.io/... is rate-limited per service hostname. Limits are not per client IP or per credential: every caller of that hostname shares one bucket.
| Default | |
|---|---|
| Burst | 1000 requests |
| Sustained refill | 300 requests per minute |
After a burst, new requests are admitted as tokens refill (on average five per second). A client that stays at or below 300 requests/minute after the burst will not be limited. A client that keeps sending faster than the refill rate will receive 429 until it slows down.
When a request is limited
The gateway returns:
| Value | |
|---|---|
| HTTP status | 429 |
X-Xg3-Status | RateLimited |
Retry-After | Whole seconds until another single request may be allowed (not the time to refill the full burst). At the default refill rate this is typically 1. |
X-Xg3-Request-Id | Correlation id (also in the JSON body) |
Content-Type | application/vnd.xg3.error+json |
json
{
"error": {
"code": "XG3_RATE_LIMIT_EXCEEDED",
"message": "Rate limit exceeded",
"request_id": "01J…"
}
}Honor Retry-After before sending another request. Do not busy-loop. Retry-After is delay-seconds (RFC 7231), minimum 1. It does not mean the burst of 1000 has been restored — only that one more request may succeed.
How to build against this
- Prefer staying under 300 requests/minute per service hostname in steady state.
- Short bursts up to 1000 are allowed; after that, wait for refill.
- On 429, wait
Retry-Afterseconds, then retry that one request. If you still receive 429, wait again using the new header. - Spread load across different services if you need higher aggregate throughput — each hostname has its own bucket.
Enrollment, certificate, and OAuth calls on the regional apex (https://{region}.xg3.io) use separate limits — see Platform API rate limits.
Parallel downloads and response pacing
All relayed requests share one tunnel. xgress3 paces each response independently: a slow or large download should not block other requests on the same service from starting or receiving their response headers.
What you should observe
- Several downloads or API calls in parallel complete without one large file stalling small requests for many seconds in the browser (“Waiting for server response”).
- Cancelling or closing a large download mid-stream should not permanently reduce how many concurrent requests the agent can handle.
If you still see long delays on small requests while a large download runs
- Confirm the agent is on a current release.
- Check the agent is connected in the console and the service is enabled.
- Review agent logs under
{data-dir}/logs/— the agent may have received requests while the client still waited at the gateway. - Contact support with
X-Xg3-Request-Idfrom a slow request.
Ingress errors you may see
These are customer-visible outcomes.
| HTTP | X-Xg3-Status or body | Typical cause |
|---|---|---|
| 429 | RateLimited | Too many requests to this service — see Service ingress rate limits; honor Retry-After |
| 502 | Policy / validation codes | Auth deny, headers too large, relay error — see Deny codes |
| 502 | Agent error JSON | Agent concurrency exceeded or backend forward failed |
| 504 | AGENT_UNAVAILABLE | No active tunnel (agent offline or reconnecting) |
The gateway validates caller identity and service policy before relaying. The agent does not validate JWT or service keys.
Platform API rate limits
Enrollment and token endpoints on the regional apex (https://{region}.xg3.io) are rate-limited separately from service ingress. If you automate enrollment or token minting, honor Retry-After on 429 (typically 60 seconds on these endpoints) and avoid tight polling loops.
| Endpoint | Guidance |
|---|---|
POST /enroll | One enrollment flow per agent; do not burst |
GET /enroll/status | Poll at reasonable intervals (agent handles enrollment internally in normal installs) |
POST /oauth/token | Cache tokens until near expiry; do not mint on every API call |
Production agents use the tunnel for renewal — not repeated CSR REST calls.
Request logging limits
When Request logging is enabled on a service, Request Lookup retains metadata for 7 days.
When detailed request logging is included in your plan and you enable body capture, xgress3 stores request and backend response bodies up to 1 MB each per request; larger bodies are truncated. Request-body capture applies only when the request is relayed to your backend.
See Request logging.
Tunnel keepalive and reconnection
- The agent sends application heartbeats on the tunnel so the connection stays active during idle periods.
- The gateway and network stack use transport keepalive to detect abrupt disconnects (for example unplugged cable) within tens of seconds.
- On tunnel loss, the agent reconnects with exponential backoff (1 s → 120 s cap, resetting to 1 s after a prior successful connection). No operator action is required.
Details: Agent operations — Reconnection. Planned regional updates use the same reconnect path — Cloud and regions — Planned maintenance.
What is not limited on the agent side
- Backend HTTP timeout: the agent does not impose a fixed timeout on calls to your backend; long-running backend work continues until the client or gateway ends the relay.
- Request body size: xgress3 streams request bodies end-to-end and does not impose a platform body-size ceiling. Large uploads are limited only by your backend (and any proxy in front of it), not by a fixed megabyte cap at the tenant hostname.
Ensure your backends and any upstream proxies tolerate the concurrency you configure on the agent.
Related documentation
- Agent operations — concurrency, reconnection, renewal, exit codes
- Ingress authentication — JWT and service keys
- Deny codes — 502 policy and validation errors; RateLimited is 429
- Troubleshooting — console workflows
- Request logging — capture and lookup