Railways
TOO MUCH DATA, THEN NOT ENOUGH
Designing alert triage for locomotive fleet monitoring — and the over-correction in the middle.
Year :
2024
Industry :
Railways
Client :
Indian Railways — Electric Locomotive Division
Project Duration :
6 months

Overview
Locomotive health data used to be trapped onboard. You found out a component was failing when the train came back to the shed, or when it didn’t.
The predictive maintenance system changed that — sensors streaming continuously, machine learning models flagging failures before they happened. Which created a new problem, and it was mine to solve: the system could now tell operators far more than they could act on.
My first design made that worse. My second made it worse in the opposite direction.
Role: Design Lead — research, IA, interface design, validation. Delivered with InovarTech.

Context
IRPMS — a predictive maintenance system built by InovarTech for Indian Railways’ electric locomotive fleet, deployed in the South Central Zone. IoT sensors on locomotives, ML models predicting component failure, and the monitoring interfaces I designed on top.
The scale sets the problem: across the network, monitoring of this kind generates roughly 14 million sensor events daily. No control room reads 14 million of anything. Everything below is about what an operator sees instead.
Users: control room operators monitoring fleet health, field engineers diagnosing on site, maintenance planners scheduling work, and division leadership tracking availability.
Research
Two weeks of contextual inquiry in locomotive control rooms and maintenance sheds. 25+ interviews with operators, technicians and engineers across divisions. Shadowing maintenance teams through both scheduled work and emergency repairs — the emergency shifts mattered most, because that’s when the interface would be under real load.
Four findings shaped everything after.
Alert fatigue was already advanced. The existing systems produced thousands of low-priority alerts a day. Operators had learned to ignore them, and that learned behaviour didn’t discriminate — critical alerts got ignored too. The problem wasn’t that operators missed alerts. It’s that they’d been trained by their own tools not to look.
Roles needed different granularity, not different volumes. Leadership wanted availability KPIs. Operators wanted current status. Technicians wanted raw diagnostic detail. Giving everyone the same view filtered differently doesn’t solve this — the data itself needed to be layered.
Field and control room are different products. Engineers diagnosing a locomotive on track need something a control room dashboard can’t be. Different device, different posture, different network.
Transparency was a precondition for trust. Operators only acted on a prediction when they could see the reasoning behind it. A black-box confidence score failed outright — not “performed worse,” failed.
The Iteration That Taught Me the Most
Version 1 — too much. I designed for the technicians. Raw sensor data surfaced, detailed metrics available, the full picture visible. Operators were overwhelmed. They’d told me they were drowning in noise, and I had handed them more, better-organised noise. The lesson I took: users need operational context, not raw metrics.
Version 2 — too little. So I stripped it back to clean status indicators. Green, amber, red. Fleet overview at a glance. This tested well with operators — and broke the technicians, who could no longer get at the diagnostic detail they needed to actually troubleshoot a fault. I had solved the first problem by creating its mirror image.
The real lesson was in the gap between those two versions: I’d been treating “how much information” as a single dial with a correct setting. It isn’t. The right amount of information is a function of who’s asking and what they’re doing right now.
Final — layered. Overview for monitoring, detail on demand for diagnosis. Progressive disclosure not as a design principle applied top-down, but as the only structure that survived contact with both user groups.

What Shipped
Fleet health overview. Geographic visualisation of locomotive status, colour-coded, with role-based views so each user lands on their own granularity rather than filtering into it.
Predictive timeline. Forecast of failure probability over time, with confidence intervals shown rather than hidden — the direct consequence of the transparency finding. An operator who can see the model’s uncertainty can decide how much weight to give it.
Three-tier alerting with auto-resolution. Green/amber/red against established industrial conventions, and — more importantly — alerts that clear themselves when the underlying issue resolves. Alert fatigue is caused as much by stale alerts as by new ones.
Mobile field interface. Offline-capable for remote track, 44px+ touch targets for gloved operation, diagnostic data and maintenance history on site. Not a responsive version of the dashboard — a different tool for a different job.
Results
The IRPMS programme’s published outcomes, as reported by InovarTech:
15% increase in locomotive uptime
20% reduction in maintenance costs
10% reduction in safety-related incidents
5% improvement in overall operational efficiency
These are programme-level results across sensors, models, integration and interface — not attributable to design alone. I’d rather point you at the verifiable number for the whole system than invent a number for my slice of it.
What I can speak to directly is what changed in the work: fleet status moved out of spreadsheets and into a live view, and diagnostic context that used to require switching between systems arrived in one place.
What I’d Do Differently
I’d have tested version 1 with both user groups before building it. The over-correction in version 2 happened because I validated the fix against the group who complained loudest, not against everyone the change touched. That’s a research design error, and it cost a full iteration cycle.
I’d instrument the alert-dismissal behaviour from day one. The single most valuable signal in a system like this is which alerts operators ignore. We could have been learning that continuously and feeding it back into prioritisation. We weren’t.
More Projects
Railways
TOO MUCH DATA, THEN NOT ENOUGH
Designing alert triage for locomotive fleet monitoring — and the over-correction in the middle.
Year :
2024
Industry :
Railways
Client :
Indian Railways — Electric Locomotive Division
Project Duration :
6 months

Overview
Locomotive health data used to be trapped onboard. You found out a component was failing when the train came back to the shed, or when it didn’t.
The predictive maintenance system changed that — sensors streaming continuously, machine learning models flagging failures before they happened. Which created a new problem, and it was mine to solve: the system could now tell operators far more than they could act on.
My first design made that worse. My second made it worse in the opposite direction.
Role: Design Lead — research, IA, interface design, validation. Delivered with InovarTech.

Context
IRPMS — a predictive maintenance system built by InovarTech for Indian Railways’ electric locomotive fleet, deployed in the South Central Zone. IoT sensors on locomotives, ML models predicting component failure, and the monitoring interfaces I designed on top.
The scale sets the problem: across the network, monitoring of this kind generates roughly 14 million sensor events daily. No control room reads 14 million of anything. Everything below is about what an operator sees instead.
Users: control room operators monitoring fleet health, field engineers diagnosing on site, maintenance planners scheduling work, and division leadership tracking availability.
Research
Two weeks of contextual inquiry in locomotive control rooms and maintenance sheds. 25+ interviews with operators, technicians and engineers across divisions. Shadowing maintenance teams through both scheduled work and emergency repairs — the emergency shifts mattered most, because that’s when the interface would be under real load.
Four findings shaped everything after.
Alert fatigue was already advanced. The existing systems produced thousands of low-priority alerts a day. Operators had learned to ignore them, and that learned behaviour didn’t discriminate — critical alerts got ignored too. The problem wasn’t that operators missed alerts. It’s that they’d been trained by their own tools not to look.
Roles needed different granularity, not different volumes. Leadership wanted availability KPIs. Operators wanted current status. Technicians wanted raw diagnostic detail. Giving everyone the same view filtered differently doesn’t solve this — the data itself needed to be layered.
Field and control room are different products. Engineers diagnosing a locomotive on track need something a control room dashboard can’t be. Different device, different posture, different network.
Transparency was a precondition for trust. Operators only acted on a prediction when they could see the reasoning behind it. A black-box confidence score failed outright — not “performed worse,” failed.
The Iteration That Taught Me the Most
Version 1 — too much. I designed for the technicians. Raw sensor data surfaced, detailed metrics available, the full picture visible. Operators were overwhelmed. They’d told me they were drowning in noise, and I had handed them more, better-organised noise. The lesson I took: users need operational context, not raw metrics.
Version 2 — too little. So I stripped it back to clean status indicators. Green, amber, red. Fleet overview at a glance. This tested well with operators — and broke the technicians, who could no longer get at the diagnostic detail they needed to actually troubleshoot a fault. I had solved the first problem by creating its mirror image.
The real lesson was in the gap between those two versions: I’d been treating “how much information” as a single dial with a correct setting. It isn’t. The right amount of information is a function of who’s asking and what they’re doing right now.
Final — layered. Overview for monitoring, detail on demand for diagnosis. Progressive disclosure not as a design principle applied top-down, but as the only structure that survived contact with both user groups.

What Shipped
Fleet health overview. Geographic visualisation of locomotive status, colour-coded, with role-based views so each user lands on their own granularity rather than filtering into it.
Predictive timeline. Forecast of failure probability over time, with confidence intervals shown rather than hidden — the direct consequence of the transparency finding. An operator who can see the model’s uncertainty can decide how much weight to give it.
Three-tier alerting with auto-resolution. Green/amber/red against established industrial conventions, and — more importantly — alerts that clear themselves when the underlying issue resolves. Alert fatigue is caused as much by stale alerts as by new ones.
Mobile field interface. Offline-capable for remote track, 44px+ touch targets for gloved operation, diagnostic data and maintenance history on site. Not a responsive version of the dashboard — a different tool for a different job.
Results
The IRPMS programme’s published outcomes, as reported by InovarTech:
15% increase in locomotive uptime
20% reduction in maintenance costs
10% reduction in safety-related incidents
5% improvement in overall operational efficiency
These are programme-level results across sensors, models, integration and interface — not attributable to design alone. I’d rather point you at the verifiable number for the whole system than invent a number for my slice of it.
What I can speak to directly is what changed in the work: fleet status moved out of spreadsheets and into a live view, and diagnostic context that used to require switching between systems arrived in one place.
What I’d Do Differently
I’d have tested version 1 with both user groups before building it. The over-correction in version 2 happened because I validated the fix against the group who complained loudest, not against everyone the change touched. That’s a research design error, and it cost a full iteration cycle.
I’d instrument the alert-dismissal behaviour from day one. The single most valuable signal in a system like this is which alerts operators ignore. We could have been learning that continuously and feeding it back into prioritisation. We weren’t.
More Projects
Railways
TOO MUCH DATA, THEN NOT ENOUGH
Designing alert triage for locomotive fleet monitoring — and the over-correction in the middle.
Year :
2024
Industry :
Railways
Client :
Indian Railways — Electric Locomotive Division
Project Duration :
6 months

Overview
Locomotive health data used to be trapped onboard. You found out a component was failing when the train came back to the shed, or when it didn’t.
The predictive maintenance system changed that — sensors streaming continuously, machine learning models flagging failures before they happened. Which created a new problem, and it was mine to solve: the system could now tell operators far more than they could act on.
My first design made that worse. My second made it worse in the opposite direction.
Role: Design Lead — research, IA, interface design, validation. Delivered with InovarTech.

Context
IRPMS — a predictive maintenance system built by InovarTech for Indian Railways’ electric locomotive fleet, deployed in the South Central Zone. IoT sensors on locomotives, ML models predicting component failure, and the monitoring interfaces I designed on top.
The scale sets the problem: across the network, monitoring of this kind generates roughly 14 million sensor events daily. No control room reads 14 million of anything. Everything below is about what an operator sees instead.
Users: control room operators monitoring fleet health, field engineers diagnosing on site, maintenance planners scheduling work, and division leadership tracking availability.
Research
Two weeks of contextual inquiry in locomotive control rooms and maintenance sheds. 25+ interviews with operators, technicians and engineers across divisions. Shadowing maintenance teams through both scheduled work and emergency repairs — the emergency shifts mattered most, because that’s when the interface would be under real load.
Four findings shaped everything after.
Alert fatigue was already advanced. The existing systems produced thousands of low-priority alerts a day. Operators had learned to ignore them, and that learned behaviour didn’t discriminate — critical alerts got ignored too. The problem wasn’t that operators missed alerts. It’s that they’d been trained by their own tools not to look.
Roles needed different granularity, not different volumes. Leadership wanted availability KPIs. Operators wanted current status. Technicians wanted raw diagnostic detail. Giving everyone the same view filtered differently doesn’t solve this — the data itself needed to be layered.
Field and control room are different products. Engineers diagnosing a locomotive on track need something a control room dashboard can’t be. Different device, different posture, different network.
Transparency was a precondition for trust. Operators only acted on a prediction when they could see the reasoning behind it. A black-box confidence score failed outright — not “performed worse,” failed.
The Iteration That Taught Me the Most
Version 1 — too much. I designed for the technicians. Raw sensor data surfaced, detailed metrics available, the full picture visible. Operators were overwhelmed. They’d told me they were drowning in noise, and I had handed them more, better-organised noise. The lesson I took: users need operational context, not raw metrics.
Version 2 — too little. So I stripped it back to clean status indicators. Green, amber, red. Fleet overview at a glance. This tested well with operators — and broke the technicians, who could no longer get at the diagnostic detail they needed to actually troubleshoot a fault. I had solved the first problem by creating its mirror image.
The real lesson was in the gap between those two versions: I’d been treating “how much information” as a single dial with a correct setting. It isn’t. The right amount of information is a function of who’s asking and what they’re doing right now.
Final — layered. Overview for monitoring, detail on demand for diagnosis. Progressive disclosure not as a design principle applied top-down, but as the only structure that survived contact with both user groups.

What Shipped
Fleet health overview. Geographic visualisation of locomotive status, colour-coded, with role-based views so each user lands on their own granularity rather than filtering into it.
Predictive timeline. Forecast of failure probability over time, with confidence intervals shown rather than hidden — the direct consequence of the transparency finding. An operator who can see the model’s uncertainty can decide how much weight to give it.
Three-tier alerting with auto-resolution. Green/amber/red against established industrial conventions, and — more importantly — alerts that clear themselves when the underlying issue resolves. Alert fatigue is caused as much by stale alerts as by new ones.
Mobile field interface. Offline-capable for remote track, 44px+ touch targets for gloved operation, diagnostic data and maintenance history on site. Not a responsive version of the dashboard — a different tool for a different job.
Results
The IRPMS programme’s published outcomes, as reported by InovarTech:
15% increase in locomotive uptime
20% reduction in maintenance costs
10% reduction in safety-related incidents
5% improvement in overall operational efficiency
These are programme-level results across sensors, models, integration and interface — not attributable to design alone. I’d rather point you at the verifiable number for the whole system than invent a number for my slice of it.
What I can speak to directly is what changed in the work: fleet status moved out of spreadsheets and into a live view, and diagnostic context that used to require switching between systems arrived in one place.
What I’d Do Differently
I’d have tested version 1 with both user groups before building it. The over-correction in version 2 happened because I validated the fix against the group who complained loudest, not against everyone the change touched. That’s a research design error, and it cost a full iteration cycle.
I’d instrument the alert-dismissal behaviour from day one. The single most valuable signal in a system like this is which alerts operators ignore. We could have been learning that continuously and feeding it back into prioritisation. We weren’t.





