Gaurav Tiwari on Keeping Trading Platforms Consistent at Scale

Institutional trading platforms depend on dozens of interconnected systems that must process the same market events in the same way, while engineers keep changing them. As those systems spread across data centres, cloud environments and locations, keeping them in step has become an architectural problem in its own right.

Gaurav Tiwari

Gaurav Tiwari, CFA, FRM, is vice president, digital assets, and leads trading technology at a global investment management firm, where he directs work on low-latency trading infrastructure, distributed event processing, observability and AI-powered engineering automation. In written answers to The Fintech Times, he set out the engineering principles he applies across nearly two decades of building platforms for large investment firms.

“A trading platform proves its value under adverse conditions, not during routine operation,” Tiwari said. “Speed is important, but predictable behaviour under stress is what allows an organisation to continue trading confidently during periods of high market volatility or infrastructure disruption.”

Part of that resilience, he said, is graceful degradation: when a platform comes under strain, non-critical services such as reporting or analytics should absorb the impact while execution and risk management carry on. “I’ve seen peripheral system failures cascade into core trading functions because those boundaries weren’t clearly established, so problems that should have remained isolated ended up disrupting critical trading systems.”

One view of the truth

The harder problem is consistency. Execution engines, risk systems, order management and post-trade processing may now run across on-premises infrastructure, cloud environments and different locations. “Even small propagation delays can leave one system acting on information another has not yet received,” Tiwari said. “In an environment where decisions are measured in microseconds, those inconsistencies directly affect risk calculations, P&L reporting, and regulatory compliance, so every system has to interpret trading activity the same way.”

His answer has been deterministic event sequencing, which gives every system the same ordered record of what happened. He found that the same design also solved a second problem. “Running multiple instances of an order management system eliminates single points of failure, but it also creates the risk of processing the same event more than once,” he said. “A sequencing layer can receive messages from every instance, remove duplicates, and emit a single ordered stream, giving every downstream system the same view of what happened.”

Sequencing alone is not enough, he added. Three other decisions matter as much. The first is a common data model. “An order, a position, or a trade should mean exactly the same thing to the execution engine, the risk platform, and every downstream system,” he said. The second is state machine replication, so that every application recognises each entity’s state transitions in the same order. “Otherwise, sooner or later, one system will process an invalid state.” The third is observability across the platform as a whole, through standardised logging, metrics and telemetry, because many problems only show up when systems are viewed together.

Choosing what to add

Every new technology layer, in Tiwari’s view, solves one problem and introduces another. “The question isn’t whether a technology is good; it’s whether it solves the problem better than the alternatives,” he said.

He gave the example of a low-latency sequencing platform his team kept on-premises because it needed microsecond-level end-to-end latency, giving up the elasticity and managed services of the cloud. “For that application, the latency requirements drove the decision. In another environment, the balance could easily shift the other way. I’ve learned to be cautious about adopting a technology simply because it’s popular.”

To bring senior stakeholders with him on such calls, he uses MoSCoW analysis, sorting requirements into Must Have, Should Have and Could Have before comparing options. When his team chose a messaging technology for the same platform, some candidates were easier to run and others delivered the performance needed at the cost of more operational complexity. “People didn’t necessarily agree on every detail, but they understood why the recommendation made sense for that particular system.”

The habit goes back to early in his career, when he designed a data warehouse for a large asset manager. His first instinct was to adopt a distributed processing framework that was then attracting attention, but the users needed fast, reliable querying more than large-scale processing. The team chose a columnar analytical database and added distributed processing later as data volumes grew. “I’ve learned to start with the business problem and let the technology follow.”

Where LLMs fit

Tiwari has also built AI-powered automation for engineering workflows. He sees the clearest value for large language models in three places: summarising the research, regulatory filings and internal documentation financial institutions produce; letting users query trusted internal data in natural language rather than writing complex queries; and helping new analysts and portfolio managers learn internal processes and tools, which reduces the support burden on engineering and operations teams.

“The biggest caution is accuracy,” he said. “LLMs still generate incorrect answers with confidence, and in financial services even a small mistake can have significant consequences. Bias, data privacy, and the protection of sensitive information all have to be built into the implementation from the beginning.”

Changing a live system

Observability should be designed in from the start, he argued. “Once a system is in production, most of the work shifts from writing software to operating, supporting, and maintaining it. If observability isn’t built into the platform from the start, engineers spend far more time trying to understand what’s happening when something goes wrong.”

For introducing major change into a live trading environment, he relies on three approaches depending on the risk involved: feature flags, so new functionality can be switched off at once without rolling back a whole deployment; running new and existing systems in parallel on the same production traffic, with a reconciliation layer comparing results in real time until the new system has proved itself across enough volume and market conditions; and, for lower-risk workflows, routing a small share of production traffic to the new system and widening it gradually.

When weighing performance against resilience, automation and maintainability, he tries to make each trade-off concrete. “Instead of debating abstract ideas, I ask questions like: Will users notice an additional hundred milliseconds of latency? Can this workflow tolerate any data loss, or is that a hard requirement?” Benchmarks, load testing and production metrics inform those answers, and documentation records why a decision was taken for the engineers who inherit it.

Asked which of AI, cloud-native infrastructure and automated trading workflows will shape trading infrastructure most, Tiwari declined to pick one. “I’ve worked through several generations of technology, and every few years our industry embraces another wave of platforms, tools, or architectural approaches,” he said. “One thing I’ve learned is that building reliable trading infrastructure still comes down to understanding the problem you’re trying to solve before deciding which technologies belong in the solution.”

The post Gaurav Tiwari on Keeping Trading Platforms Consistent at Scale appeared first on The Fintech Times.

Read More

Leave a Reply

Your email address will not be published. Required fields are marked *