The most useful diagnostic tools in a computer owner’s free toolkit in 2026 are MemTest86+ or PassMark MemTest86 for RAM, CrystalDiskInfo for drive health, OCCT (plus GPU-Z sensors) for GPU and VRAM faults, and HWiNFO64 rail readings confirmed with a multimeter for the PSU — Windows’ built-in utilities are a good first pass, but none of them can actually prove hardware guilt. Below is the exact tool per component, plus what a pass and a fail actually look like.
- Quick map: symptom to tool to verdict
- Built-in Windows diagnostics: what they catch and miss
- RAM faults: MemTest86+ and TestMem5
- Drive health: read SMART in CrystalDiskInfo, not chkdsk
- GPU faults: OCCT and FurMark 2, decoded with GPU-Z
- PSU: the component software can’t fully test
- The order that saves the most time
- FAQ
Quick map: symptom to tool to verdict
| Symptom | Built-in first check | Best free tool | Result that confirms a hardware fault |
|---|---|---|---|
| Random BSODs, app crashes, corrupt files | Windows Memory Diagnostic (mdsched) | MemTest86+ / MemTest86 | Any error count above zero |
| Drive vanishes, stutters, slow boots | chkdsk /scan | CrystalDiskInfo | Non-zero Reallocated or Pending Sectors |
| Crashes or artifacts only under GPU load | Reliability Monitor history | OCCT VRAM test / memtest_vulkan | VRAM errors or artifacts at stock clocks |
| Instant black-screen restarts in games | Event Viewer (Kernel-Power 41) | OCCT Power test + HWiNFO64 | +12 V rail below 11.4 V under load |
Built-in Windows diagnostics: what they catch and miss
- Reliability Monitor and Event Viewer: the best starting point for pattern-finding. Kernel-Power Event ID 41 means the machine lost power or hard-locked; WHEA-Logger errors usually implicate the CPU or unstable voltages; repeated Display Event 4101 (“driver stopped responding”) points at the GPU.
- Windows Memory Diagnostic (mdsched.exe): catches a grossly dead DIMM in about 15–20 minutes, but runs far fewer test patterns than MemTest86 — a clean result does not clear marginal or overclocked RAM.
- chkdsk /scan, sfc /scannow, DISM: repair file-system and OS corruption. Useful, but they operate on software — they say nothing about whether the drive itself is dying.
- powercfg /batteryreport: on laptops, compares design capacity to current full-charge capacity; below roughly 70% of design means a worn battery, not a Windows problem.
- Driver Verifier: deliberately stresses drivers to expose bad ones. Powerful, but it can boot-loop a broken system — set a restore point first and only use it after hardware has passed.
The pattern: built-in tools are excellent at ruling software in or out. Proof of failing hardware needs the utilities below.
RAM faults: MemTest86+ and TestMem5
Boot MemTest86+ (open source) or PassMark’s MemTest86 from a USB stick and run the full four-pass cycle — expect several hours for 32 GB of DDR5, so overnight runs are normal. Reading the result is simple: the error counter should be zero. A single error means failure. Errors repeating in the same address range indicate a physically bad stick; test one DIMM at a time to identify which.
The critical nuance: run the test twice, once with XMP/EXPO enabled and once at JEDEC defaults. Errors only with the profile enabled mean an unstable memory overclock or a marginal integrated memory controller, not a dead module — drop the frequency one step or add roughly 0.05 V to DRAM voltage and retest. For fine-tuning overclocks inside Windows, TestMem5 with the anta777 Extreme profile or HCI MemTest surface marginal instability faster than MemTest86’s patterns.
Drive health: read SMART in CrystalDiskInfo, not chkdsk
CrystalDiskInfo translates raw SMART data into a verdict. These are the attributes that matter and their real thresholds:
| Attribute | ID | Meaning | Action threshold |
|---|---|---|---|
| Reallocated Sector Count | 05 | Bad sectors remapped to spare area | Raw value above 0 — back up immediately |
| Current Pending Sector Count | C5 | Sectors the drive can’t read, awaiting remap | Above 0 — drive is failing |
| Uncorrectable Sector Count | C6 | Reads that failed even after retries | Above 0 — replace the drive |
| UltraDMA CRC Error Count | C7 | Data corrupted between board and drive | Any count — replace the SATA cable first, not the drive |
| NVMe Media & Data Integrity Errors | — | Flash-level read/write failures | Rising count — start a warranty claim |
| NVMe Percentage Used | — | Wear against rated endurance | 100% or more — past rated life |
For deeper repair and secure-erase functions, use the vendor’s own free tool: Samsung Magician, Western Digital Dashboard, or Seagate SeaTools. CrystalDiskMark then gives a performance sanity check: expect roughly 500–550 MB/s sequential reads on SATA, 3,000–3,500 on PCIe 3.0 NVMe, and 5,000–7,000 on Gen4. If a Gen4 drive benches near 1,700 MB/s, check CrystalDiskInfo’s Transfer Mode line — a drive linked at PCIe 3.0 or x2 lanes is a slot or lane problem, not a dying SSD.
GPU faults: OCCT and FurMark 2, decoded with GPU-Z
Run OCCT’s 3D and VRAM tests (the free license caps each run at one hour, which is enough). Any VRAM error above zero at stock clocks means failing memory or an over-ambitious factory OC — retest with a small negative core-clock offset to separate the two. The standalone memtest_vulkan utility does the same VRAM check without OCCT’s overhead.
FurMark 2 is the thermal worst case; read it through GPU-Z’s sensors. A Hot Spot delta beyond roughly 25–30 °C over the core temperature means the cooler isn’t seating properly (dried paste or a loose mount). Memory junction temperatures past about 105 °C on GDDR6X cards explain thermal shutdowns under load. Interpretation matters: visible artifacts at stock clocks equal hardware failure, while crashes that began right after a driver update — with clean sensors — are software. Check the minidump in NirSoft BlueScreenView or WhoCrashed; a crash chain ending in nvlddmkm.sys or atikmdag.sys confirms the graphics stack.
PSU: the component software can’t fully test
No utility measures a PSU’s real output — motherboard sensors are approximate at best. OCCT’s Power test helps by loading the CPU and GPU simultaneously to provoke a marginal unit, and HWiNFO64 shows the rails. ATX spec allows ±5% on the main rails:
| Rail | Nominal | In-spec range | Suspect below |
|---|---|---|---|
| +12 V | 12.00 V | 11.40–12.60 V | 11.4 V under load |
| +5 V | 5.00 V | 4.75–5.25 V | 4.75 V |
| +3.3 V | 3.30 V | 3.135–3.465 V | 3.13 V |
If software reports an out-of-spec rail, verify with a multimeter on a PSU connector or a basic PSU tester (usually $15–$30) before buying a replacement — sensor readings can simply be wrong.
For the classic “dies only in games” complaint, do the sizing math: sustained draw is roughly CPU boost power + GPU board power + ~70 W for the board, drives and fans. Example: a 150 W CPU plus a 220 W GPU plus 70 W is about 440 W sustained. PSUs are happiest below ~70% of rated capacity, so that build wants a 650–750 W quality unit. It gets worse on older PSUs: modern graphics cards produce transient spikes near double their rated draw for microseconds — ATX 3.x supplies are built to absorb them, while older ATX 2.x units trip over-current protection and black-screen the system. A tired 500 W PSU powering that 440 W build is a textbook cause.
The order that saves the most time
- First ten minutes: Reliability Monitor pattern check plus a CrystalDiskInfo SMART scan — both are instant and rule out the two most common culprits.
- Free fixes next: reseat cables (non-zero C7 counts), roll back a driver if crash history lines up with an update, and clear any overclocks including XMP/EXPO.
- Long test second: MemTest86 four passes before blaming anything else for random BSODs — RAM is the most frequent silent offender.
- Load tests last: OCCT’s GPU and Power tests once RAM and storage are confirmed clean, since a mid-stress crash muddies the evidence.
FAQ
Can stress tests damage my hardware?
Power-virus tests like FurMark 2 run components at their thermal ceiling, but modern silicon throttles itself to stay safe. Still, watch temperatures and stop the test if a GPU hotspot passes ~110 °C or the CPU sits pinned at its thermal limit.
Windows Memory Diagnostic found nothing — is my RAM fine?
Not proven. Its pattern set is far lighter than MemTest86’s. Run four full passes of MemTest86+ before clearing the memory — and if XMP/EXPO is on, test with the profile both on and off.
Do I need every tool listed here?
No. CrystalDiskInfo, HWiNFO64, and a MemTest86 USB stick cover the majority of real-world faults. Add OCCT and memtest_vulkan only when a crash pattern already points at the GPU or PSU.






