Cover image

Count the Edges Before You Count the GPUs

TL;DR for operators A graph-transformer job that is too large or too slow for one GPU presents two separate questions: how to divide the work, and how many GPUs are worth allocating. Treating the second question as “more is faster” is unreliable. In Scalable and Adaptive Parallel Training of Graph Transformer on Large Graphs, Lin, Madduri, and Kandemir1 report 8-GPU A100 speedups of 6.1x for ogbn-proteins, 3.3x for ogbn-products, and 4.2x for Reddit. The same accelerator count therefore produces materially different returns across graph workloads. The preferred parallelization strategy also changes across graph and hardware configurations. ...

September 6, 2026 · 7 min · Zelina
Cover image

When Trains Meet Snowstorms: Turning Weather Chaos into Predictable Rail Operations

A delayed train is easy to complain about and surprisingly hard to explain. The passenger sees one number: five minutes late, twelve minutes late, cancelled, chaos. The operator sees a messier object. Was the train already late when it entered the station? Did the station itself add delay? Was the delay caused by snow, low visibility, wind, passenger boarding, a single-track bottleneck, equipment failure, or simply the accumulated sins of every previous station on the route? ...

January 26, 2026 · 20 min · Zelina
Cover image

Flip the Switch: How Heterogeneous Agents Learn to Restore the Grid

A power outage is not one problem. It is a queue of smaller, uglier problems pretending to be one. Which switches can be closed? Which loads should come back first? Which distributed generators are available? Which lines will overheat if a local microgrid gets too ambitious? Which voltage limits will quietly make the elegant restoration plan unusable? In a control room, these questions arrive together, under time pressure, with the usual helpful accompaniment of incomplete information and operational consequences. ...

November 20, 2025 · 15 min · Zelina
Cover image

Bridges and Biases: How LLMs Are Learning to Inspect Infrastructure

TL;DR for operators Bridge teams do not usually lack data. They lack enough expert time to turn dense inspection data into clear, defensible decisions. That is the operational gap this paper tries to narrow: not by replacing bridge engineers with a chatbot in a hard hat, thankfully, but by using multimodal LLMs to translate non-destructive evaluation contour maps into structured condition assessments and maintenance recommendations.1 ...

July 21, 2025 · 16 min · Zelina