Debugging your mining rig (hardware & HiveOS software)
A rig that crashes, a graphics card that disappears, a miner that keeps restarting: these are the most common failures in GPU mining, and they can cost you hours. In this 2021 video (in French), we show you how to track down the cause methodically, first on the hardware side, then on the software side with the HiveOS tools: dashboard, miner logs, watchdog and command line.
1. Hardware: isolate the problem
Our example is a rig with 8 non-LHR RTX 3070 cards, powered by two power supplies. The right reflex: don't examine everything at once, reduce the number of suspects instead. Unplug the second power supply and run the rig on the 4 cards of the first one. If it stays stable, the problem lies with the other 4 cards; if it still crashes, it is on the first line.
Then check the most fragile parts:
- The riser LEDs: they show the 12 V and 3.3 V power. The blue and red LEDs must be lit.
- The risers themselves: they are consumables costing €5 to €7. When a card causes trouble, don't hesitate to replace its riser.
- The connections: every cable and every connector must be properly seated.
2. When the hardware is not to blame: overclocking
One night, we spent three hours on a new rig that would start mining and then crash. We unplugged one power supply to keep only 4 cards: the problem remained. We replaced the motherboard, then the risers: still the same issue. The graphics cards were brand new. The hardware was not to blame: the problem came from the overclocking.
In our experience, when a rig mines for a while and then crashes, it is often an overclocking setting that is too aggressive. In this case the culprit was the memory: on the same GPU, here the RTX 3070, the onboard memory can come from different brands. We were used to non-LHR cards with Samsung memory, which takes a memory overclock of at least 2400, or even 2600 to 2700. On LHR models, manufacturers often fit another memory, which we consider lower quality, and which is set between 1400 and 1700, 1800 at most. A setting designed for Samsung memory therefore makes these cards crash.
3. Monitor your rig in HiveOS
On the worker page, several indicators help you spot a problem:
- Temperatures: on NVIDIA cards, HiveOS only shows the GPU temperature, not the memory temperature, even though the memory is often what causes crashes. AMD cards do report the memory temperature. The RTX 3080 and 3090, which use GDDR6X, often have memory heat issues.
- Messages: one area collects configuration changes, warnings and errors. Check it regularly.
- The card table: temperatures, fan speeds and power draw of each GPU.
- The load average: it measures the CPU load. On our rig it is 0.28 over the last minute and 0.16 on average over the last 15 minutes, which is normal. It should not exceed the number of CPU cores multiplied by 2.
To dig deeper, the Miners menu, then Action and Miner log, shows the output of the mining software. That is where you see the errors that come before a crash.
4. Automate restarts with the watchdog
The watchdog, "the dog that keeps watch", automatically restarts the miner or the rig when a condition is met. For example, you can:
- restart the miner after three minutes if it detects a problem;
- reboot the rig after a set number of minutes;
- reboot if the load average goes above a threshold;
- monitor the hashrate: if T-Rex drops below 250 MH/s, something is wrong, and the watchdog restarts the miner.
We had a rig that crashed every 5, 10 or 15 hours, and for a few weeks we could not find the cause. With the watchdog, the rig restarted on its own, which limited the mining time lost while we searched. The cause turned out to be the memory of the RTX 3080-type cards, which ran too hot and made the miner crash: a classic.
5. The command line for in-depth diagnosis
In HiveOS, Remote access then Hive Shell Start opens a command line, as if you were plugged directly into the rig. A few useful commands:
- net-test: checks the connection to the HiveOS servers and that there is no proxy problem;
- miner log: shows the miner output, including the percentage of valid shares;
- nvidia-info: lists the NVIDIA GPUs, numbered from 0;
- gpu-fans-find followed by the card number: spins that card's fans so you can physically identify it in the rig;
- top: shows CPU and memory usage per process (Shift + P to sort by CPU, Shift + M by memory).
Finally, keep an eye on the shares: the percentage of valid shares should stay around 98 to 99%. Below that, a card has a real problem, either overclocking or hardware.
What about today?
Since Ethereum moved to proof of stake in September 2022, Ethereum can no longer be mined with graphics cards, and GPU rigs have become marginal. The method still applies to any mining machine: isolate the parts one by one, monitor temperatures and logs, and automate restarts while you look for the cause.
To set up HiveOS properly, see our HiveOS tutorial (worker, overclocking, flight sheet) and find all our videos on the tutorials page. Moving to ASIC miners? Our team can advise you: contact us.

