BlogArticle

FluxRPC Engineering

Solana RPC Failover: Why Your Primary and Backup Need Different Failure Domains

FluxRPC Team
10 min read

A production Solana application should not depend on a single RPC path.

But having two RPC providers does not automatically give you real failover.

If your primary and secondary RPC share the infrastructure, network dependencies or serving architecture that caused the failure, both can become unavailable at the same time.

That happened in August 2026 when a Teraswitch network-layer failure disrupted Solana RPC connectivity across multiple providers and regions. Solana dApps reported that trading APIs had no healthy RPC path left to fail over to, despite routing through multiple providers.

The lesson is simple:

A real Solana RPC backup needs a different failure path, not just a different provider name.

FluxRPC is designed around that principle. Its RPC serving and state-processing architecture is separated from the full validator stack, giving applications a materially different path between Solana state and their application.

And that different path does not require an enterprise contract. The current FluxRPC Developer plan is $38/month for 250 GB of bandwidth and 100 requests per second, using the same FluxRPC infrastructure as the larger plans.

What is Solana RPC failover?

Solana RPC failover is the ability for an application to continue accessing current Solana state when its normal RPC path becomes unavailable, stale or degraded.

A typical setup has:

Primary RPC provider

and

Secondary RPC provider

But the important part is not having two URLs.

The important part is ensuring that the second path can remain healthy when the first one fails.

That means evaluating failure domains.

A failure domain is any dependency capable of affecting multiple parts of your infrastructure at the same time.

For Solana RPC, that can include:

  • network providers
  • hosting infrastructure
  • geographic locations
  • validator infrastructure
  • validator software
  • state-processing architecture
  • RPC serving infrastructure
  • upstream data sources
  • control planes

If your primary and secondary RPC share the dependency that fails, your failover plan can disappear exactly when you need it.

What did the Teraswitch incident show?

On August 12, 2026, a routing failure at Teraswitch knocked roughly 90 Solana validators offline for approximately 33 minutes.

About 28.83% of staked SOL became delinquent, although Solana continued finalizing blocks and recovered without a finality halt.

For RPC users, the more important lesson happened above the consensus layer.

Jupiter reported that its Ultra, Swap V2 order and Trigger APIs began returning 5xx errors and timeouts.

Its postmortem explains why:

a network-layer failure took out RPC connectivity across multiple providers and regions

The result was that Jupiter had no healthy RPC path available to fail over to for the affected APIs.

That is a real-world example of correlated RPC failure.

Jupiter had multiple providers.

What it temporarily did not have was a surviving failure path.

Why can a primary and backup RPC fail together?

Because vendor diversity is not the same thing as infrastructure diversity.

Imagine this:

Primary RPC: Provider A

Secondary RPC: Provider B

From your application, that looks redundant.

But underneath, both providers might share some combination of:

  • the same hosting provider
  • the same network provider
  • similar validator infrastructure
  • the same geographic dependency
  • the same type of validator-coupled serving architecture

When the shared dependency fails, the fact that the API keys came from two different companies does not help.

Your users do not care that two vendors failed independently.

They care that your application stopped working.

For a trading platform, wallet, exchange, payment application or liquidation system, that can mean:

  • trades cannot be built
  • transactions cannot be simulated
  • fresh blockhashes cannot be obtained
  • balances or positions appear stale
  • liquidations stop
  • transaction status disappears
  • users retry actions
  • support volume increases
  • confidence in the application falls

Infrastructure failures quickly become product failures.

And product failures become trust problems.

A backup that fails the same way is not a backup

The usual question is:

“Who is our secondary RPC provider?”

The better question is:

“What can break our primary that does not also break our secondary?”

That changes how you evaluate RPC redundancy.

You are no longer comparing just:

  • latency
  • price
  • RPS
  • method support

You are also comparing:

  • failure architecture
  • data paths
  • state-processing paths
  • serving infrastructure
  • geographic distribution
  • operational dependencies

This is where FluxRPC becomes particularly relevant.

Why does FluxRPC have a different failure path?

FluxRPC does not simply expose another conventional validator-backed RPC endpoint.

The architecture is described in Why FluxRPC.

FluxRPC separates RPC serving and state processing from the full validator stack.

At a simplified level, a conventional validator-coupled RPC path can look like:

Solana network → full validator stack → RPC serving → application

FluxRPC instead separates the stages:

Solana network data → FluxRPC ingestion → state processing → RPC serving → application

That distinction matters.

The complete validator stack does not have to sit directly in the serving path for every RPC request.

FluxRPC can ingest the Solana data it needs, reconstruct and maintain the required state, and then serve RPC workloads through separate serving infrastructure.

That creates a different failure boundary.

FluxRPC is not independent of Solana. It still relies on Solana as the blockchain and underlying data source.

The important distinction is that RPC serving is separated from the full validator stack.

That is a much more useful property in a failover provider than simply giving the same architecture another hostname.

How does FluxRPC isolate failures internally?

The architecture also separates several pieces of the RPC path instead of treating everything as one monolithic validator process.

FluxRPC uses separate components for:

  • data ingestion
  • state processing
  • data shards
  • RPC serving
  • geographic delivery

This means failures can potentially be isolated at different layers rather than propagating through one combined validator and RPC process.

FluxRPC also uses multiple data sources and verification mechanisms to protect against stale state.

For streaming, the Yellowstone implementation documents a multi-producer architecture where multiple Solana nodes can feed the gRPC service at the same time. If one producer stops providing data, another can continue supplying the stream.

Again, this does not mean FluxRPC cannot fail.

It means it has been designed with different and separated failure domains.

Why does that matter if FluxRPC is your primary RPC?

The same architecture that makes FluxRPC useful as a secondary path also matters if FluxRPC is your primary provider.

A good primary RPC should be designed to contain failures before they reach your application.

Separating state processing, data ingestion and serving infrastructure can reduce the number of cases where one internal failure automatically becomes a full RPC outage.

So this is not a pitch to use FluxRPC only as a backup.

FluxRPC can be:

  • your primary RPC
  • your secondary RPC
  • part of an active-active RPC architecture

The failover point is simply one of the clearest examples of why its architecture matters.

Why does that matter if FluxRPC is your secondary RPC?

If you are already happy with your current RPC provider, you do not need to replace it to benefit from FluxRPC.

The architecture can be:

Application → Primary provider

with:

Application → FluxRPC

as the alternative path.

The commercial question then becomes very simple:

What does it cost to add a materially different RPC failure path to your production stack?

For many teams, the answer starts at $38 per month.

What does a different RPC failure path cost?

FluxRPC offers pay-as-you-go bandwidth at $6 per 100 GB.

This creates a useful way to think about failover cost.

For $38 per month, a team can maintain a production-capable FluxRPC path with:

  • 250 GB of included bandwidth
  • 100 requests per second
  • Yellowstone access
  • WebSocket access
  • the same underlying FluxRPC infrastructure used across the plans

That is not the cost of replacing your entire RPC stack.

It is the cost of adding a different path into it.

For a production application with real users, that is a relatively small infrastructure cost compared with discovering during an outage that both RPC providers depended on the same failure domain.

How should you test Solana RPC failover?

Adding the second endpoint is only step one.

The second path needs to be continuously usable.

A production system should test at least:

  • Can the endpoint respond?
  • Is its Solana slot advancing?
  • Is its block height current?
  • Can it return a fresh blockhash?
  • Do the RPC methods your application actually needs work?
  • Is latency still inside your operational threshold?
  • Can transactions still be submitted if your application needs that path?
  • Are WebSocket or Yellowstone streams still progressing?

See the full FluxRPC RPC documentation.

Why HTTP 200 is not enough

A dangerous RPC failure is not necessarily a dead endpoint.

The endpoint may still return HTTP 200 while serving stale state.

For example:

Primary RPC: slot 500,000,000

Secondary RPC: slot 499,999,600

Both endpoints respond.

Only one is current enough to use.

A useful health check therefore needs to evaluate chain freshness, not just server availability.

Your failover rule should be closer to:

fail over if the primary is unreachable, too slow, no longer advancing, too far behind the network, or failing critical methods

rather than:

fail over only when the server stops returning HTTP responses

What should trigger a failover?

There is no universal threshold because applications have different requirements.

A wallet may tolerate conditions that a liquidation engine cannot.

But production systems should explicitly decide what constitutes an unhealthy RPC.

Possible signals include:

  • repeated transport errors
  • excessive request latency
  • slot progression stopping
  • unacceptable slot lag
  • block height not advancing
  • getLatestBlockhash failing
  • transaction submission failing
  • WebSocket subscriptions no longer progressing
  • Yellowstone streams becoming stale

Failover should be an application decision based on the things the application actually depends on.

Why reliability becomes a user trust issue

Infrastructure teams naturally think about RPC outages in technical terms.

Users do not.

A user sees:

“My trade failed.”

“My transaction is stuck.”

“My balance is wrong.”

“The app is down.”

They do not see:

“Two nominally independent RPC providers happened to share a correlated infrastructure dependency.”

That distinction matters when designing production systems.

The purpose of failover is not simply to improve an uptime statistic.

It is to prevent an infrastructure dependency from becoming a visible product failure.

The Teraswitch lesson for every Solana application

The August incident provides a very clean answer to the question:

Can multiple Solana RPC providers fail at the same time?

Yes.

Jupiter reported exactly that.

A network-layer incident affected RPC connectivity across multiple providers and regions, leaving some of its APIs without a healthy RPC route.

The correct response is not to conclude that multi-provider redundancy does not work.

It is to build better multi-provider redundancy.

That means deliberately selecting providers with different failure characteristics.

Why FluxRPC belongs in that architecture

FluxRPC provides standard Solana RPC infrastructure through FluxRPC:

FluxRPC's differentiator in a failover architecture is not simply that it is another Solana RPC provider.

It is that its serving architecture is separated from the full validator stack, with separate data ingestion, state processing and serving layers.

That makes it useful as a primary RPC.

It also makes it unusually useful as the different path in a multi-provider architecture.

And that different path can start at $38 per month.

For Consideration

If your application depends on Solana RPC, you should ask two questions:

Do we have more than one RPC provider?

and then:

Can those providers fail for the same reason?

The second question is more important.

The Teraswitch incident showed that multiple providers and multiple regions can still converge on the same infrastructure failure.

FluxRPC approaches that problem from the architecture itself.

Its RPC serving and state processing are separated from the full validator stack, creating a different operational path from conventional validator-coupled RPC infrastructure.

Use FluxRPC as your primary if that architecture makes sense for your application.

Use it alongside your existing provider if you want failure-domain diversity.

But do not confuse two endpoints with two failure paths.

For production infrastructure, the difference can cost as little as $38 per month.

And you usually discover whether that difference mattered on the worst possible day.