Jason M. Oliverjmoliver.ai
I solve infrastructure problems that live between ownership boundaries.

Work Examples

These are high-level examples of the kinds of infrastructure problems I work through. Details are intentionally general so customer, employer, and internal system information stays protected.

Distributed storage performance that did not have one obvious cause

Some storage problems do not start as a storage problem. A customer or internal team may see slow access, uneven throughput, or application impact while the platform itself appears healthy from a dashboard view.

My work in that kind of situation is to trace the path instead of arguing from one layer. I look at client behavior, protocol access, metadata pressure, network or fabric behavior, storage telemetry, timing, and workload shape until the likely constraint becomes testable.

The useful output is usually a clearer timeline, a smaller failure domain, and a set of validation steps that support, engineering, and the customer can act on.

AI/HPC workload readiness before the environment is trusted at scale

GPU or compute availability does not prove a workload is ready. AI and HPC environments can underperform because of data movement, file or object access patterns, scheduler placement, checkpoint paths, fabric behavior, or assumptions hidden between teams.

I focus on the surrounding platform: where the data comes from, how it moves, how the workload is placed, what telemetry can prove, and where latency or throughput pressure can appear before it becomes a production incident.

That kind of review is most useful before a team adds more hardware, changes a storage design, or declares the environment ready based on partial signals.

Virtualization and infrastructure escalations where every team has a partial view

In large enterprise environments, an issue can cross VMware, storage, network, operating system, and application boundaries. Each team may have valid evidence, but no single layer explains the whole behavior.

My role is to slow the investigation down just enough to make it useful: define what is known, separate symptoms from causes, compare telemetry to user impact, and build a testable explanation that narrows the next step.

This is where years of escalation work help. The goal is not to win a theory; it is to get the issue into a form that can be reproduced, remediated, or handed to the right engineering owner with useful evidence.

Portable diagnostics and runbooks that survive handoff

Good troubleshooting should leave something behind. In customer-facing and internal engineering work, I have often needed evidence collection to be repeatable, low-risk, and understandable by the next person who picks up the case.

That can mean scripts, packet captures, structured logs, health checks, validation notes, or a short runbook that explains what was checked and why it mattered.

The point is practical: reduce repeated manual work, make escalation evidence cleaner, and help teams avoid rediscovering the same problem under a different name.

Short Field Notes

Optical transceiver instability

Intermittent path issues can look like storage or host instability until link behavior, thermal context, and hardware evidence are compared together.

High-density link and port telemetry

Healthy-looking links can still produce intermittent errors. Host-side and switch-side evidence often need to be compared before the pattern is clear.

Authentication path latency

Not every storage complaint is storage. External identity systems, network hops, and access patterns can materially change perceived platform performance.