TrackMyGPUBack to Workspace →

GPU Doctor

Get your GPU working.

Check your setup, understand a failure, and know what to try next.

Check a workload

What are you checking?

Use peak workload memory, including batches and context—not just model weights. Multi-GPU memory is not automatically pooled.

Check a running GPU

In your GPU's SSH terminal, download and inspect the probe, then run it using your workload's Python environment from its data directory.

Download diagnostic probe
python doctor_probe.py

The probe checks packages, free disk space, GPU access, a tiny GPU operation, and default localhost ports. It installs nothing and sends no data. Port checks do not verify application health.

Explain an error log

When you check, the report and error lines are sent to TrackMyGPU for analysis. Doctor doesn't save them. Known credential patterns are removed, but review your input for other sensitive data.

Your diagnosis

+

A clearer next step.

Enter what you know. Doctor will separate verified checks, likely issues, and information still needed.

Doctor suggests fixes. It never installs packages, changes files, restarts a GPU, or launches compute.