E-Series Performance Analyzer v3.5.1
EPA v3.5.1
E-Series Performance Analyzer aka EPA v3.5.1 was released today.
- Unresolved E-Series system failure events are now exported as Prometheus alets
- InfluxDB was updated to latest and greatest 1.12.2
- Minor optimization of docker-compose.yml
Unresolved or "active" system failures were already gathered in 3.5.0 and before, but they weren't exported as Prometheus metrics. Now they are.
As an example, here's the usual "LUN not on preferred path":
# HELP eseries_active_failures_total Number of active failures
# TYPE eseries_active_failures_total gauge
eseries_active_failures_total{failure_type="nonPreferredPath",object_ref="1",object_type="notOnPreferredPath",sys_id="600A098000F63714000000005E79B17B",sys_name="EF570"} 1.0
Prometheus users can now use Prometheus Alert Manager to scrape these in addition to whatever other stuff Prometheus metrics already offer.
An even better alerting feature would be to look at various object metrics, but that can be done client-side. It's challenging to do it server-side (in Collector and Prometheus exporter) because without access to different hardware it is impossible to know what conditions should constitute "best practice" alerts. What the above gives the user is the reasonable and obvious: if SANtricity says it's a failure, it gets exported.
InfluxDB v1.12.2 released weeks ago doesn't have any features or fixes that benefit EPA 3 (that I know of), but it's generally helpful to have the latest version.
Next steps
The InfluxDB v1 updates (two so far this year) show EPA 3 is very much a viable option.
The Grafana component it comes with is old (v8), but I assume most users use own Grafana. The only reason I don't want to upgrade that Grafana is I would also have to test the dashboards, which I am not a big fan of. (If anyone wants to do it and submit a pull request, please!)
I'm satisfied with where EPA 3 and ESC 4 are now, especially when compared with other current E-Series-focused alternatives.