EV Charger Root Cause Analysis and Field Failure Engineering
Most field failure investigations stop at the layer the investigator owns. That is why the same fault keeps returning.
A session that fails intermittently can originate in the contactor, the firmware state machine, the protocol exchange or the backend, and each specialist can demonstrate their layer is behaving correctly. All of them can be right while the fault persists, because the fault lives in the interaction.
Send us a failure caseFor CPO, OEM, operations lead
Preserve the evidence first
The most common reason an investigation cannot conclude is that the unit was disturbed before anyone looked. Power cycling clears volatile state. Disassembly destroys thermal and mechanical evidence. Replacing the suspect part removes the only sample.
- Capture logs and backend records before touching the unit
- Photograph the installation as found, including cable routing and enclosure condition
- Record firmware version, configuration and recent update history
- Retain the failed part rather than scrapping it
- Note the environmental and supply conditions at the time of failure
Work across the layers, not within one
The investigation moves between hardware, firmware, protocol and backend rather than settling in one. A contactor that fails to close may be a hardware fault, a firmware timing assumption, or a remote command the backend believed it sent. Distinguishing them requires the message trace and the hardware, together.
Reproduce before concluding
A hypothesis that cannot be reproduced is a theory. Reproduction on the bench, under the conditions the field unit met, is what separates a root cause from a plausible explanation. Intermittent faults often need the specific combination of temperature, supply condition and timing that the site provided.
Corrective action, then verification
A change believed to fix the problem is not a fix until subsequent field data confirms it. Corrective actions are closed against evidence rather than against the belief that the change was correct, which means the loop stays open for a defined period after deployment.
Containment while you investigate
Investigation takes time and the fleet stays in service. Containment decisions, whether that is a configuration change, a firmware hold, a service instruction or a targeted replacement, are made early and separately from the root cause work.
A failure mode that passes every acceptance test
The hardest failures to find are the ones that were not present when the unit shipped. They pass end-of-line test, install cleanly, work for months, and then appear in a car park a long way from anyone who could have caught them.
Internal cabling is the example we know best because we found it in our own returns rather than in any standard. Cable quality and the way cable is arranged inside the enclosure produce real field failures. Nothing about that is visible at acceptance, which is exactly why it survives into the field and why it took return analysis rather than inspection to identify.
That is what root cause work is for. Not to explain a unit that failed, but to find the pattern across units that failed for reasons the test plan did not anticipate, and then to change the test plan. Our field failure rate of roughly 0.5 percent is the accumulated result of that loop, not of any single control.
Want this applied to your own site?
Send us a failure caseTechnically reviewed by Anees P K, Director of Technology. Last reviewed 2026-08-29.
Frequently asked questions
We already replaced the part and it failed again. Can you still investigate?
Yes, though it is harder. The replaced part was the primary evidence. Backend records, firmware history and the installation itself remain useful.
Can you investigate chargers you did not build?
Yes. Cross-layer analysis is a method rather than a product feature, though access to firmware behaviour and backend records determines how far it can go.
How do you tell a hardware fault from a firmware one?
By correlating the message trace with the hardware behaviour at the same timestamps. Each in isolation usually looks correct, which is why single-layer investigations stall.
What should our field team do differently?
Capture logs and photograph the installation before touching anything, record firmware version and update history, and retain failed parts rather than scrapping them.
When is a corrective action closed?
When field data after deployment confirms the failure mode has stopped, not when the change is shipped.
Send us a failure case
Tell us the symptom, the firmware version and what evidence still exists. We will tell you what can be concluded from it.
Send us a failure case