Posted on

Mitsubishi FR-F840 E.UVT Undervoltage Fault Troubleshooting Guide: A Complete Diagnostic Approach from DC Bus Charging to Voltage Detection Circuit Failure

Introduction

In industrial automation systems, variable frequency drives (VFDs) are widely used for controlling motors in applications such as pumps, fans, compressors, conveyors, machine tools, and production equipment. As the operating environment becomes more demanding, VFDs are exposed to voltage fluctuations, temperature stress, dust contamination, and long-term electrical aging.

When a drive displays a fault code, many technicians immediately associate the alarm name with the failure location. For example, an overcurrent alarm is often considered an IGBT failure, while an undervoltage alarm is assumed to be caused by insufficient input voltage.

However, modern industrial VFDs use complex protection and monitoring systems. A fault code usually represents the condition detected by the control system, not necessarily the exact failed component.

A typical example is the Mitsubishi FR-F840-00620-2-60, a 400V-class high-power inverter with approximately 62A rated output current and 30kW motor capacity. In field repair work, this model may experience the E.UVT (Undervoltage Trip) alarm. Especially when the drive reports E.UVT immediately after power-on, the actual cause is often not simply an external power supply problem.

Possible causes include:

  • Three-phase input abnormality;
  • Rectifier circuit failure;
  • DC bus charging problems;
  • Pre-charge resistor or relay failure;
  • DC bus voltage detection circuit malfunction;
  • Control power instability;
  • Main control board sampling errors.

This article provides a systematic troubleshooting method for Mitsubishi FR-F840 E.UVT faults, especially for cases where the inverter alarms immediately after energization and the power module has already been confirmed to be normal.


Mitsubishi FR-F840 inverter E.UVT undervoltage fault troubleshooting with DC bus voltage measurement using a digital multimeter

1. Basic Structure of Mitsubishi FR-F840-00620-2-60

Before troubleshooting an undervoltage fault, it is necessary to understand the internal power structure of the inverter.

The Mitsubishi FR-F840 belongs to the F800 series, designed mainly for industrial fan, pump, HVAC, and general-purpose drive applications.

The basic energy conversion process is:

Three-phase AC input

R / S / T

↓

Input protection and filtering

↓

Rectifier bridge

↓

DC bus

P(+) / N(-)

↓

DC capacitors

↓

IGBT inverter module

↓

U / V / W output

↓

Motor

The main functions of each section are:

Rectifier section

Converts three-phase AC voltage into DC voltage.

DC bus section

Stores energy and stabilizes the DC voltage.

IGBT inverter section

Converts DC voltage into variable-frequency AC output for motor control.

Control board

Responsible for:

  • Voltage monitoring;
  • Current detection;
  • Protection logic;
  • PWM generation;
  • Fault judgment.

The E.UVT fault is mainly related to the section:

AC input
↓
Rectification
↓
DC bus charging
↓
DC voltage detection

Mitsubishi FR-F840 variable frequency drive repair showing pre-charge circuit, DC bus section, and voltage detection control board inspection during E.UVT fault diagnosis

2. What Does E.UVT Mean?

E.UVT stands for:

Undervoltage Trip

It means:

The inverter has detected that the DC bus voltage has dropped below the allowable operating threshold and has activated protection.

For a 400V-class inverter:

The approximate DC bus voltage can be calculated as:

DC voltage ≈ AC voltage × 1.414

For example:

380VAC × 1.414 ≈ 537VDC

Normally, the FR-F840 DC bus voltage should be approximately:

500–560VDC

depending on the actual input voltage.

If the DC bus voltage decreases significantly, for example:

300VDC

or lower, the control system determines that the inverter cannot safely drive the IGBT section and triggers:

E.UVT

3. Why Is “Immediate E.UVT After Power-On” Important?

The timing of the fault provides valuable diagnostic information.

There are three typical situations:


Case 1: E.UVT During Normal Operation

Example:

The inverter runs normally for several minutes and then suddenly trips.

Possible causes:

  • Utility voltage fluctuation;
  • Insufficient transformer capacity;
  • Large load startup on the same power network;
  • Loose input contactor;
  • Poor cable connection.

Case 2: E.UVT During Acceleration

Possible causes:

  • Excessive motor load;
  • Acceleration time too short;
  • DC regenerative energy problems;
  • Weak power supply.

Case 3: E.UVT Immediately After Power-On

This is the most important condition.

At this moment:

  • The motor has not started;
  • The IGBT output is inactive;
  • Load influence is minimal.

Therefore, the problem is usually located in the power supply establishment and voltage detection circuits.

The main inspection areas are:

  1. AC input;
  2. Rectifier bridge;
  3. Pre-charge circuit;
  4. DC bus capacitors;
  5. DC voltage detection circuit.

4. Step-by-Step Troubleshooting Procedure

Step 1: Check Three-Phase Input Voltage

First measure:

R-S
S-T
T-R

Normal values:

380–440VAC

The three phases should be balanced.

Example:

Normal:

R-S = 402V
S-T = 401V
T-R = 403V

Abnormal:

R-S = 400V
S-T = 395V
T-R = 250V

Possible causes:

  • Phase loss;
  • Damaged contactor;
  • Loose terminal;
  • Power supply problem.

5. Step 2: Measure DC Bus Voltage

This is the most important measurement.

Measure between:

P(+)

and

N(-)

Expected value:

Approximately:

500–560VDC

The result determines the troubleshooting direction.


Situation A: DC Bus Voltage Does Not Build Up

Example:

P-N = 50VDC

The inverter cannot establish the DC bus.

Possible causes:


1. Rectifier Bridge Failure

The IGBT module may be normal, but the rectifier section can still fail.

Possible faults:

  • Open rectifier diode;
  • Damaged rectifier module;
  • Input phase failure;
  • Internal connection problem.

Important:

A normal IGBT does not mean the complete power section is normal.

The rectifier and inverter sections are independent.


2. Pre-Charge Circuit Failure

Large-capacity inverters cannot directly charge large DC capacitors because the initial charging current would be extremely high.

Therefore, they use a pre-charge circuit:

AC input

↓

Rectifier bridge

↓

Pre-charge resistor

↓

DC capacitors

↓

Bypass relay/contactor

↓

Normal operation

If any of the following fail:

  • Pre-charge resistor open;
  • Relay does not activate;
  • Relay contact burned;
  • Drive circuit failure;

the DC bus cannot charge correctly, resulting in E.UVT.


Situation B: DC Bus Voltage Is Normal but E.UVT Still Appears

This situation is very common during professional repairs.

Example measurement:

P-N = 530VDC

but the inverter still displays:

E.UVT

This means:

The actual DC voltage is normal, but the control system believes the voltage is too low.

The suspected area is:

DC bus voltage detection circuit.


6. DC Bus Voltage Detection Principle

The CPU cannot directly measure 500VDC.

Therefore, the inverter uses a voltage detection circuit:

DC 500V

↓

High-voltage resistor divider

↓

Isolation circuit

↓

ADC sampling

↓

CPU calculation

↓

Protection judgment

If this circuit fails:

Actual voltage:

530VDC

Detection result:

200VDC

The CPU will incorrectly trigger:

E.UVT

Common Detection Circuit Failures

1. High-Voltage Resistor Drift

High-voltage resistors operate continuously under electrical stress.

After years of operation:

  • Resistance increases;
  • Resistance decreases;
  • Internal cracks occur.

The voltage division ratio changes, causing incorrect measurement.


2. Optocoupler Aging

Some inverter designs use isolation components.

After long operation:

  • Optical transmission efficiency decreases;
  • Signal amplitude becomes incorrect.

3. Detection IC Failure

Possible problems:

  • ADC input abnormality;
  • Operational amplifier damage;
  • Reference voltage failure.

7. Control Power Supply Problems

The control board requires stable low-voltage supplies.

Important rails include:

+5V Power Supply

Used by:

  • CPU;
  • Digital circuits;
  • Memory.

If:

5V drops to 4.5V

the CPU may misjudge voltage signals.


+15V Power Supply

Used for:

  • Gate drive circuits;
  • Analog detection circuits.

+24V Power Supply

Used for:

  • Relays;
  • External control interfaces.

Unstable control power can cause:

  • E.UVT;
  • CPU errors;
  • Communication faults.

8. Common Repair Mistakes

Mistake 1: Replacing the IGBT Immediately

Many technicians see a power-related alarm and replace the IGBT module.

This is often unnecessary.

The IGBT may be completely normal.


Mistake 2: Only Measuring Input Voltage

Checking:

R/S/T voltage normal

does not prove the inverter is healthy.

The technician must also check:

P-N DC bus voltage

because the failure may exist in:

  • Rectification;
  • Pre-charge;
  • Voltage detection.

Mistake 3: Assuming a Normal Power Module Means the Main Circuit Is Good

The main power system includes:

  • Rectifier;
  • DC capacitors;
  • Pre-charge circuit;
  • Voltage detection;
  • IGBT inverter.

All sections must be verified.


9. Example Repair Case: FR-F840-00620-2-60 Immediate E.UVT

Equipment

Model:

Mitsubishi FR-F840-00620-2-60

Power:

30kW

Fault:

E.UVT immediately after power-on

Inspection Process

Step 1

Three-phase input voltage checked.

Result:

Normal.

External power supply was excluded.


Step 2

IGBT module checked.

Result:

Normal.

Power module failure was excluded.


Step 3

DC bus measured.

Result:

Approximately:

530VDC

The DC bus was successfully established.


Step 4

Voltage detection circuit inspected.

Finding:

The DC voltage feedback signal was abnormal.

Cause:

High-voltage divider components had drifted from their original values.

After repairing the detection circuit:

The inverter returned to normal operation.


This case demonstrates an important principle:

An undervoltage alarm does not always mean the actual voltage is low.

The failure may exist in the measurement system.


10. Recommended Diagnostic Strategy for High-Power VFD Repair

For inverters above 30kW, technicians should follow a fixed troubleshooting sequence:

Confirm fault code

↓

Analyze fault timing

↓

Measure AC input

↓

Measure DC bus voltage

↓

Check charging circuit

↓

Check voltage detection

↓

Check control power supply

↓

Repair and test

Avoid unnecessary replacement of expensive components such as:

  • IGBT modules;
  • Control boards;
  • Main boards.

The correct approach is:

Measure first, diagnose second, replace components last.


Conclusion

For Mitsubishi FR-F840-00620-2-60 inverters displaying E.UVT undervoltage faults, especially when the alarm occurs immediately after power-on, the troubleshooting focus should not be limited to the external power supply.

The correct diagnostic sequence is:

  1. Verify three-phase input voltage;
  2. Measure DC bus voltage;
  3. Check rectifier and pre-charge circuits;
  4. Verify DC voltage feedback detection;
  5. Check control board power supplies.

When the power module has already been confirmed normal, technicians should pay special attention to:

  • DC bus voltage detection circuits;
  • Pre-charge charging circuits;
  • Control board sampling circuits.

The key principle of industrial inverter troubleshooting is:

A fault code shows what the drive detected, not necessarily where the failure occurred.

Only by combining electrical measurements, circuit understanding, and systematic analysis can technicians accurately locate faults, reduce unnecessary component replacement, and improve repair efficiency.

Posted on

Siemens SINAMICS S120 F30025 and F00004 Fault Analysis: Why an Instant Overtemperature Alarm on a Cold Drive Usually Indicates a Motor Module Problem

Introduction

The Siemens SINAMICS S120 drive system is widely used in high-performance industrial applications such as CNC machine tools, robotics, packaging systems, printing equipment, metallurgy production lines, and complex automation platforms. Compared with standard frequency converters, the SINAMICS S120 adopts a highly modular architecture, providing excellent flexibility and performance.

However, the advanced structure of the S120 also makes troubleshooting more complex. Many engineers tend to interpret fault messages literally. When they see:

  • F30025 – Power unit overtemperature
  • F00004 – Drive overtemperature

their first assumption is usually:

  • The cabinet temperature is too high.
  • The cooling fan has failed.
  • The heat sink is overheating.
  • The drive has insufficient ventilation.

In many real-world repair cases, however, a very different situation occurs:

The drive reports an overtemperature fault immediately after power-up, even though the Motor Module is completely cold and has not had any opportunity to generate heat.

This situation appears contradictory:

  • The system reports an overheating condition.
  • The drive temperature is normal.
  • The cooling system is working.
  • The parameter temperature reading looks normal.
  • The fault cannot be reset.

The actual problem is often not a real thermal overload, but rather:

An abnormal temperature detection circuit, internal power unit fault, or Motor Module hardware failure causing the SINAMICS protection system to trigger an overtemperature alarm.

This article explains the fault mechanism, diagnostic process, differences between CU320 and Motor Module faults, and practical troubleshooting methods for SINAMICS S120 F30025/F00004 alarms.


Close-up view of Siemens SINAMICS S120 Motor Modules inside an industrial control cabinet showing F30025 overtemperature fault indication and drive status LEDs during troubleshooting.

1. Understanding the SINAMICS S120 System Architecture

Before analyzing F30025, it is important to understand the hardware structure of SINAMICS S120.

Unlike a conventional inverter where the control and power section are integrated into one unit, the SINAMICS S120 uses a modular design.

A typical structure is:

              PLC / Controller
                    |
                    |
                CU320-2 PN
             (Control Unit)
                    |
                DRIVE-CLiQ
                    |
        ---------------------------
        |                         |
   Motor Module              Motor Module
        |                         |
      Motor 1                 Motor 2

The two main parts are:


1.1 CU320 Control Unit

The CU320-2 PN is responsible for:

  • Motion control
  • Parameter management
  • Communication
  • DRIVE-CLiQ network management
  • Topology identification
  • Fault processing

A typical model:

6SL3040-1MA01-0AA0
SINAMICS CU320-2 PN

The CU320 can be considered the “brain” of the SINAMICS system.

However, it does not directly drive the motor power stage.


1.2 Motor Module / Power Unit

The Motor Module performs the actual high-power functions:

  • IGBT switching
  • Motor current control
  • Voltage measurement
  • Current measurement
  • Power semiconductor protection
  • Temperature monitoring

Most F30025 faults originate from this section.

This means:

The fault is displayed by the CU320, but the actual fault source is usually inside the Motor Module.


Siemens SINAMICS S120 CU320-2 PN control unit and Motor Modules installed in an industrial electrical cabinet with DRIVE-CLiQ connections and power module wiring.

2. What Does SINAMICS S120 F30025 Really Mean?

F30025 – Power unit overtemperature

Siemens defines F30025 as:

Power unit: Overtemperature

Many engineers interpret this as:

“The drive temperature is physically too high.”

However, this is incomplete.

SINAMICS monitors multiple temperature points, including:

  • Heat sink temperature
  • IGBT module temperature
  • Semiconductor junction temperature
  • Internal electronic board temperature

The most important value is often:

IGBT Junction Temperature

The temperature inside an IGBT chip can be significantly higher than the temperature measured on the outside of the drive.

For example:

Heat sink surface:

40°C

while internally:

IGBT junction temperature > protection limit

This is possible because heat must travel through several thermal layers:

IGBT chip
    |
Silicon layer
    |
Solder layer
    |
Base plate
    |
Heat sink

Every layer introduces thermal resistance.

Therefore:

A normal external temperature does not always mean the internal semiconductor temperature is normal.


3. Understanding F00004 Drive Overtemperature Fault

Another common SINAMICS alarm is:

F00004 – Drive overtemperature

This fault generally relates to:

  • Power unit temperature monitoring
  • Cooling system problems
  • Temperature feedback abnormalities

When F30025 and F00004 appear together, the system believes that the power section temperature protection has been activated.

However, the key question is:

Is the drive actually hot?


4. The Most Important Diagnostic Clue: Fault Appears Immediately After Power-Up

This is the most valuable diagnostic information.

A typical real overheating fault follows this sequence:

Drive starts
      |
Motor operates
      |
Power loss generates heat
      |
Temperature rises
      |
Protection activates

This process normally takes:

  • several seconds,
  • minutes,
  • or longer depending on load.

However, if the fault appears:

  • immediately after power-on,
  • before motor operation,
  • while the cabinet is cold,

then the situation is different.

The likely sequence becomes:

Temperature detection signal abnormal
              |
              |
Controller interprets high temperature
              |
              |
F30025/F00004 protection triggered

In this situation, the system is not detecting real heat.

It is detecting an abnormal temperature signal.


5. Why Engineers Often Suspect the CU320

Because the CU320 is connected to everything, many technicians initially suspect it.

The reasoning is usually:

“The fault appears on the control panel, so the control unit must be bad.”

However, this is a common misunderstanding.

A useful comparison is a car:

If the dashboard shows:

“Engine temperature too high”

the dashboard itself is not necessarily defective.

The problem may be:

  • engine sensor,
  • wiring,
  • ECU,
  • cooling system.

SINAMICS works similarly:

Motor Module
      |
Temperature sensor
      |
Power electronics
      |
DRIVE-CLiQ
      |
CU320
      |
Fault displayed

The display location is not the same as the fault location.


6. How to Determine CU320 Fault or Motor Module Fault

Method 1: Check When the Fault Appears

Case A:

Fault appears immediately after power-up.

Most likely:

  • Motor Module temperature sensing failure
  • Power unit electronics fault
  • Internal communication problem

Case B:

Fault appears after running.

Check:

  • Cooling fan
  • Air filter
  • Cabinet temperature
  • Motor load
  • PWM frequency

7. Using Temperature Parameters for Diagnosis

SINAMICS provides temperature monitoring parameters.

One commonly checked parameter is:

r0037
Power unit temperature

Example:

Faulty drive:

30°C

Healthy drive:

30°C

This indicates:

  • Ambient temperature is normal.
  • The temperature feedback value is not obviously abnormal.

However, engineers should be careful:

A normal r0037 value does not completely eliminate a temperature sensing problem.

Because:

Different sensors may exist:

  • Heat sink temperature sensor
  • IGBT junction temperature calculation
  • Internal protection sensor

A possible situation is:

Heat sink temperature normal

        ↓

IGBT temperature sensing abnormal

        ↓

F30025 triggered

8. Why Replacing the CU320 Usually Does Not Solve F30025

The CU320 manages:

  • Control
  • Communication
  • Parameters

It does not contain:

  • IGBT modules
  • Power transistors
  • Heat sink sensors
  • Power temperature circuits

If the Motor Module has:

  • failed NTC sensor,
  • damaged temperature sampling circuit,
  • power board problem,

replacing the CU320 will not correct the fault.

The correct troubleshooting order is usually:

  1. Verify CU320 communication.
  2. Confirm fault source.
  3. Inspect Motor Module.
  4. Repair or replace the power unit.

9. Possible Internal Failures Inside the Motor Module

9.1 Temperature Sensor (NTC) Failure

The most common cause.

NTC sensors work by changing resistance according to temperature.

Typical failures:

  • Open circuit
  • Short circuit
  • Resistance drift

If the signal becomes incorrect:

The controller may interpret it as:

“Temperature too high.”


9.2 Temperature Sampling Circuit Failure

The temperature signal path may include:

  • Amplifiers
  • Filtering circuits
  • ADC inputs
  • Protection comparators

Example:

Normal:

NTC sensor
     |
Voltage change
     |
ADC conversion
     |
Controller

Fault:

ADC input abnormal

     |

Controller sees overtemperature

     |

F30025

9.3 Aging of Power Electronics

Long-term operation can cause:

  • Capacitor degradation
  • Solder fatigue
  • Thermal cycling damage
  • PCB aging

Applications with frequent acceleration/deceleration are especially demanding.

Examples:

  • CNC machines
  • Press machines
  • Hoists
  • High-speed equipment

9.4 IGBT Module Problems

IGBT aging may cause:

  • Increased conduction losses
  • Higher heat generation
  • Abnormal thermal behavior

However:

If the alarm occurs immediately after power-up, a pure IGBT overheating problem is less likely.


10. Recommended Troubleshooting Procedure

For SINAMICS S120 F30025/F00004 faults, the recommended procedure is:


Step 1: Confirm Fault Timing

Record:

  • Immediate after power-on?
  • After motor operation?
  • Under heavy load?

Step 2: Check Cooling System

Verify:

  • Cooling fans
  • Air filters
  • Cabinet ventilation
  • Ambient temperature

Step 3: Use Siemens Diagnostic Software

The most effective tools are:

  • Siemens Startdrive
  • TIA Portal
  • STARTER

Check:

  • Fault history
  • Fault value
  • Component number
  • Drive object

The most important information is:

Which drive object generated the fault?

Step 4: Compare with a Healthy Module

If multiple Motor Modules exist:

Compare:

  • Temperature values
  • Current values
  • Status information
  • Diagnostic data

Step 5: Decide Repair or Replacement

If the Motor Module is expensive:

Repair evaluation may be worthwhile.

Possible repair areas:

  • Temperature sensing circuit
  • Power interface board
  • Driver board
  • Internal electronics

11. Common Troubleshooting Mistakes

Mistake 1: Trusting the Fault Name Literally

Seeing:

“Overtemperature”

does not always mean:

“High temperature exists.”

Industrial fault messages describe the protection condition, not always the physical cause.


Mistake 2: Replacing CU320 First

Because CU320 displays the fault, technicians sometimes replace it first.

However:

For F30025:

The Motor Module should normally be investigated first.


Mistake 3: Working Without Siemens Software

The BOP display provides only limited information.

SINAMICS S120 is designed to be diagnosed through:

  • Startdrive
  • TIA Portal
  • STARTER

Without software:

Important information is hidden.


12. Practical Case Summary

A typical real case:

A Siemens SINAMICS S120 system reported:

F30025
Power unit overtemperature

F00004
Drive overtemperature

The operator checked:

  • Cabinet temperature normal
  • Cooling system normal
  • Drive was cold
  • Temperature parameter approximately 30°C

The faults appeared immediately after power-up.

Initial suspicion:

CU320 control unit.

After further diagnosis:

The problem was confirmed to be related to the Motor Module, not the CU320.

The final conclusion:

The Motor Module had an internal power unit temperature monitoring problem.


13. Final Conclusion

SINAMICS S120 F30025 and F00004 faults require careful analysis.

The most important diagnostic rule is:

The location where the fault is displayed is not necessarily the location where the fault occurs.

For a normal overheating condition:

  • Temperature rises gradually.
  • Load operation causes the fault.
  • Cooling problems are involved.

For a false overtemperature alarm:

  • Fault appears immediately.
  • Drive is still cold.
  • Temperature values appear normal.
  • Motor Module internal diagnostics become the focus.

When F30025/F00004 occurs immediately after startup, engineers should prioritize inspection of:

  • Motor Module temperature sensing circuits
  • Power unit electronics
  • Internal protection circuits
  • DRIVE-CLiQ communication

Understanding the architecture of SINAMICS S120 allows technicians to avoid unnecessary replacement of expensive components such as CU320 controllers and quickly locate the true fault source.

For high-value industrial drives, accurate diagnosis is often more valuable than simply replacing parts. A systematic troubleshooting approach can significantly reduce downtime, repair costs, and production losses.

Posted on

ABB Soft Starter Power Supply Board Failure Diagnosis and Repair Analysis: A Technical Study Based on the 1SFB536268D1007 Power Board

Abstract

ABB soft starters are widely used in industrial motor control systems for applications requiring reduced starting current, smooth acceleration, controlled stopping, and motor protection. As one of the most important internal components, the power supply board plays a critical role in converting incoming AC power into multiple stable DC voltages required by the control system, trigger circuits, communication modules, relays, and monitoring circuits.

When the power supply board fails, the soft starter may experience various problems, including internal fault alarms, failure to start, communication errors, unstable operation, or protection trips. Unlike obvious failures such as damaged thyristor modules or burned power components, power supply board failures are often hidden and require systematic troubleshooting.

This article analyzes a real-world failure case involving an ABB soft starter power supply board model 1SFB536268D1007. A batch of replacement power boards was purchased, but one unit generated a fault after installation by the customer. Through analysis of the board structure, switching power supply design, common failure mechanisms, and practical diagnostic procedures, this article provides a comprehensive troubleshooting method for industrial engineers and repair technicians.


Close-up view of an ABB soft starter power supply board with switching transformers, capacitors, connectors, and electronic components for industrial motor control system repair and troubleshooting.

1. The Importance of the Power Supply Board in ABB Soft Starters

In industrial maintenance, engineers usually focus on the main power components when a soft starter fails:

  • Thyristor modules;
  • Main control board;
  • Motor condition;
  • Three-phase input power;
  • Bypass contactor.

However, many real-world failures originate from the power supply board.

Modern ABB soft starters are not simple power switching devices. They contain a complete embedded control system consisting of:

  • Microprocessor control;
  • Current measurement circuits;
  • Trigger pulse generation;
  • Communication interfaces;
  • Protection monitoring;
  • Relay outputs;
  • User interface circuits.

The power supply board provides the necessary electrical energy for all these systems.

A simplified structure is:

Three-phase AC Input
          |
          |
  Power Semiconductor Section
          |
          |
        Motor


Control System:

Power Supply Board
          |
          +---- CPU Control
          |
          +---- Thyristor Trigger Circuit
          |
          +---- Current Detection
          |
          +---- Communication Module
          |
          +---- Relay Output

The power supply board can therefore be considered the “energy center” of the soft starter.

If any output voltage becomes unstable, the entire soft starter may malfunction even when the main power components are completely healthy.


ABB PSTX soft starter power supply board installed inside an industrial control cabinet, showing PCB components, wiring terminals, and electrical connections for motor control applications.

2. Structural Analysis of ABB 1SFB536268D1007 Power Supply Board

Based on the physical inspection of the 1SFB536268D1007 board, it adopts a typical industrial multi-output isolated switching power supply design.

The main functional sections include:

  1. AC input filtering and rectification;
  2. High-frequency switching conversion;
  3. Isolation transformer circuits;
  4. Multiple DC voltage outputs;
  5. Control interface circuits;
  6. Relay and auxiliary power circuits.

2.1 AC Input and Rectifier Section

The input section is usually located near the high-voltage side of the PCB.

Typical components include:

  • EMI filter components;
  • Surge protection devices;
  • Safety capacitors;
  • Rectifier bridge;
  • High-voltage electrolytic capacitors.

The main function is:

Convert AC input into a stable DC bus voltage.

The conversion process is:

AC Input
   |
   |
EMI Filtering
   |
   |
Rectification
   |
   |
High Voltage DC Bus

This section directly faces industrial electrical environments, including:

  • Voltage surges;
  • Lightning impulses;
  • Switching transients;
  • Grid disturbances.

Therefore, it is one of the areas most vulnerable to damage.


2.2 High-Frequency Isolation Switching Power Supply Section

The board contains several yellow magnetic components visible on the PCB.

These are not traditional low-frequency transformers. They are high-frequency isolation transformers used in switching power supplies.

Their function is:

Convert high-voltage DC into several isolated low-voltage supplies.

Typical conversion process:

High Voltage DC
        |
        |
PWM Controller
        |
        |
High Frequency Transformer
        |
        +---- +5V
        |
        +---- +12V
        |
        +---- +15V
        |
        +---- +24V

Different circuits inside the soft starter require different supply voltages.

Typical applications:

VoltageApplication
+5VCPU and logic circuits
+3.3VDigital control system
+15VTrigger and driver circuits
+24VRelays and external interfaces

If any of these voltages becomes abnormal, the soft starter may generate an internal fault.


2.3 Control Interface Section

The upper connectors and terminal blocks connect the power supply board with:

  • Main control PCB;
  • External control terminals;
  • Communication modules;
  • Feedback detection circuits.

A power supply board may appear to work normally, but abnormal interface signals can still cause:

  • Start command failure;
  • Communication errors;
  • Internal protection alarms.

3. Why One Replacement Board Failed After Installation

In practical industrial repair, it is common to encounter this situation:

“A customer receives several replacement boards. Most work normally, but one board produces a fault after installation.”

This does not necessarily indicate installation error.

Several causes should be considered.


3.1 Aging During Long-Term Storage

Electronic components can deteriorate even without operation.

The most common aging component is the electrolytic capacitor.

Electrolytic capacitors may experience:

  • Reduced capacitance;
  • Increased ESR;
  • Increased ripple current;
  • Reduced filtering performance.

For example:

Normal capacitor:

470uF
ESR = 0.05Ω

Aged capacitor:

470uF
ESR = 2Ω

A capacitance meter may still show an acceptable value, but the switching power supply may become unstable.

Typical symptoms:

  • Starts normally when cold;
  • Fault occurs after several minutes;
  • Voltage drops under load.

3.2 Switching Power Supply Feedback Failure

A switching power supply depends on a feedback loop consisting of:

  • PWM controller;
  • Optocoupler;
  • TL431 reference circuit;
  • Feedback resistors.

If the feedback loop fails:

Output voltage may increase

Example:

24V output becomes 30V.

Possible consequences:

  • Downstream IC damage;
  • Protection activation;
  • Control malfunction.

Output voltage may decrease

Example:

24V output drops to 18V.

Possible consequences:

  • Relay cannot energize;
  • CPU resets;
  • Communication failure.

3.3 Incorrect Input Voltage Version

Industrial equipment often has multiple voltage versions.

The same soft starter series may include:

  • 110VAC control supply version;
  • 230VAC version;
  • 400VAC version.

If the installed power board does not match the equipment voltage configuration, abnormal operation may occur.

For example:

Equipment designed for:

230VAC control supply

but connected to:

400VAC supply

may result in:

  • Immediate fault;
  • Component overheating;
  • Permanent damage.

Therefore, checking the equipment nameplate before replacement is essential.


3.4 The Soft Starter Itself May Cause the Fault

A common mistake during troubleshooting is assuming:

“New power board = guaranteed good.”

This is not always true.

The original soft starter may have another hidden problem:

  • Damaged thyristor module;
  • Short circuit in trigger circuit;
  • Abnormal load;
  • Damaged bypass circuit.

The new power board may simply expose the existing system problem.


4. Standard Troubleshooting Procedure for ABB Soft Starter Power Boards

Step 1: Confirm Whether the Fault Follows the Board

If multiple identical boards are available, perform a comparison test.

Example:

Soft starter A:

Install board No.1:

Fault occurs.

Install board No.2:

Works normally.

Conclusion:

Board No.1 has a high probability of failure.

If all boards fail:

The problem is likely in the soft starter system.


Step 2: Verify Input Voltage

Measure the actual voltage supplied to the power board.

Record:

Input Voltage:
Frequency:
Connection Method:

Do not only confirm that voltage exists.

Industrial electronics require:

  • Correct voltage level;
  • Stable waveform;
  • Correct frequency.

Step 3: Measure DC Output Voltages

The most important measurements are:

Test PointNormal Range
+5V4.8–5.2V
+15V14–16V
+24V22–26V

Fault examples:

5V normal, 24V abnormal

The problem is likely in the 24V secondary circuit.

All outputs abnormal

The primary switching section should be investigated.


Step 4: Check Output Ripple

Many technicians only measure DC voltage.

However, switching power supplies must also be checked for ripple.

Example:

A multimeter shows:

24V DC

but an oscilloscope reveals:

3V peak-to-peak ripple

The system may still malfunction.

Recommended oscilloscope measurements:

  • 5V ripple;
  • 15V ripple;
  • 24V ripple.

Step 5: Inspect Critical Components

Electrolytic Capacitors

Check:

  • Leakage;
  • Bulging;
  • ESR value;
  • Temperature condition.

Optocouplers

Check:

  • LED input side;
  • Output transistor switching.

PWM Controller

Check:

  • Startup voltage;
  • Switching waveform;
  • Drive signal.

Power MOSFET

Check:

  • Drain-source short circuit;
  • Gate abnormality;
  • Leakage current.

5. Common Fault Symptoms and Diagnostic Direction

Symptom 1:

No Display After Power-On

Possible causes:

  • Input fuse failure;
  • Switching power supply startup failure;
  • Primary MOSFET damage.

Inspection:

Check whether the high-voltage DC bus exists.


Symptom 2:

Display Works but Motor Cannot Start

Possible causes:

  • Insufficient 24V supply;
  • Relay failure;
  • Trigger supply abnormality.

Symptom 3:

Internal Fault Alarm

Possible causes:

  • Unstable power supply;
  • CPU supply reset;
  • Communication voltage abnormality.

Symptom 4:

Bypass Fault

Focus inspection on:

  • Relay circuit;
  • Bypass control;
  • Feedback detection.

6. Repair Recommendations for ABB Power Supply Boards

6.1 Replace Aging Capacitors Correctly

Do not select capacitors only according to capacitance.

Important specifications:

  • Voltage rating;
  • Temperature rating;
  • ESR;
  • Ripple current capability.

Industrial applications should use:

  • 105°C capacitors;
  • High-frequency switching capacitors.

6.2 Perform Load Testing After Repair

A repaired power board should not only pass no-load testing.

Recommended testing:

  • 25% load;
  • 50% load;
  • 75% load;
  • 100% load.

Monitor:

  • Output voltage stability;
  • Temperature rise;
  • Alarm conditions.

6.3 Avoid Blind Component Replacement

Replacing random ICs without measurement usually increases repair difficulty.

The correct approach is:

Fault Alarm
     |
Confirm Model
     |
Measure Input
     |
Measure Output
     |
Compare Good Board
     |
Locate Circuit Section
     |
Replace Failed Component
     |
Load Verification

7. Practical Lessons from Industrial Repair Experience

The ABB 1SFB536268D1007 power supply board failure case demonstrates an important principle:

Small electronic modules can determine the reliability of an entire industrial control system.

When dealing with replacement or refurbished boards, engineers should consider:

  • Storage aging;
  • Capacitor degradation;
  • Hidden power supply instability;
  • Compatibility issues;
  • Secondary equipment faults.

A professional repair process should include:

  1. Visual inspection;
  2. Electrical measurement;
  3. Functional comparison;
  4. Component-level diagnosis;
  5. Final operational testing.

8. Conclusion

The ABB soft starter power supply board 1SFB536268D1007 is a critical component responsible for supplying stable energy to the entire control system. Although it is not the main power switching element, its failure can completely disable the soft starter.

When a replacement board generates faults after installation, engineers should not immediately assume that the board is correct or that the customer installation is wrong. A systematic diagnostic approach is required:

  • Verify the soft starter model;
  • Confirm the alarm code;
  • Check input voltage;
  • Measure all DC outputs;
  • Analyze switching power supply circuits;
  • Inspect capacitors, feedback circuits, and switching devices;
  • Perform load testing.

Through structured troubleshooting and board-level repair methods, ABB soft starter power supply failures can be accurately identified and repaired, reducing downtime and avoiding unnecessary replacement costs in industrial automation systems.

Posted on

Deep Analysis of SINAMICS S120 F30005 and F30021 Faults After IGBT Replacement: How a -78A W-Phase Current Offset Reveals a Current Feedback Circuit Failure

Abstract

Siemens SINAMICS S120 drive systems are widely used in high-performance industrial applications, including CNC machines, robotics, printing equipment, semiconductor manufacturing, packaging systems, and other precision automation fields. Due to their advanced modular power structure, SINAMICS S120 power units have very strict requirements for power semiconductors, current measurement circuits, gate drive circuits, and protection feedback systems.

During field maintenance, one of the most challenging situations is when a SINAMICS S120 Motor Module experiences a severe overheating event caused by cabinet cooling failure. After the internal cabinet temperature rises abnormally, the drive may report faults such as:

  • F30021 – Power unit: Ground fault
  • F30005 – Power unit: Overload I²T

In many repair cases, technicians replace major power components, including:

  • IGBT power modules
  • Current transformers (CT sensors)
  • Internal fuses

However, after replacement, the drive still reports F30005 shortly after power-up. The drive may start normally when cold, but after approximately one minute the alarm appears again.

A typical example is:

  • SINAMICS S120 Motor Module
  • Model: 6SL3120-1TE21-8AA3
  • Rated output: 3AC 400V / 18A
  • DC link voltage: 600V

After IGBT and CT replacement, the diagnostic parameter shows:

Phase U offset: -0.04 A
Phase V offset: -0.27 A
Phase W offset: -78.35 A

This abnormal W-phase current offset becomes the key evidence. It indicates that the actual problem is most likely not the IGBT itself, but a failure inside the current feedback measurement circuit.

This article explains the causes, diagnostic methods, and repair strategy for SINAMICS S120 F30005/F30021 faults, using this case as a practical example.


Siemens SINAMICS S120 F30021 F30005 fault diagnosis with -78.35A W-phase current offset

1. Overview of SINAMICS S120 Motor Module Structure

1.1 Basic Power Structure

The SINAMICS S120 system is based on a modular drive architecture. A typical configuration consists of:

Three-phase AC Supply

        ↓

Line Module

        ↓

DC Link 600V

        ↓

Motor Module

        ↓

IGBT Inverter Bridge

        ↓

U/V/W Output

        ↓

Motor

Inside the Motor Module, several critical circuits work together:

  • IGBT power switching stage
  • Gate driver circuit
  • DC-link voltage monitoring
  • Three-phase current measurement
  • Temperature monitoring
  • Short-circuit protection
  • Ground fault detection

A failure in any of these circuits can generate power unit faults.


2. Understanding SINAMICS S120 Fault F30021

2.1 Meaning of F30021

Fault code:

F30021 – Power unit: Ground fault

is usually interpreted as:

The power unit has detected an abnormal leakage current or ground fault condition.

Many technicians immediately assume:

  • Motor insulation failure
  • Motor cable short circuit
  • IGBT breakdown

These are possible causes, but they are not the only causes.

The SINAMICS S120 does not simply measure insulation resistance to determine this fault. Instead, it uses:

  • Phase current feedback
  • Current vector calculation
  • Power stage protection algorithms

The drive continuously checks the relationship between the three-phase output currents:

IU + IV + IW = 0

Under normal conditions:

IU = 10A
IV = 10A
IW = 10A

The system is balanced.

However, if the current measurement circuit is incorrect:

IU = 10A
IV = 10A
IW = 80A

The controller may interpret this imbalance as abnormal leakage current and trigger F30021.

Therefore:

A false current feedback signal can also create a ground fault alarm.


SINAMICS S120 F30021 F30005 fault progression from overheating to current sensor failure

3. Understanding SINAMICS S120 Fault F30005

3.1 What Does I²T Overload Mean?

Fault:

F30005 – Power unit: Overload I²T

does not always mean that the motor is mechanically overloaded.

I²T protection is a thermal protection model.

The principle is:

Thermal stress = Current² × Time

A small current increase over a long period can accumulate enough thermal stress to trigger protection.

For example:

At 10A:

10² = 100

At 50A:

50² = 2500

The thermal effect increases dramatically.

The drive calculates the estimated thermal stress of the power module. When the calculated value exceeds the permitted limit, F30005 occurs.


SINAMICS S120 power board analysis with W-phase CT current feedback fault

4. Why Does F30005 Appear One Minute After Power-On?

This timing information is extremely important.

If the IGBT is completely shorted:

  • Fault usually appears immediately.
  • The drive trips instantly.
  • F30021 normally occurs very quickly.

However, in this case:

  • Drive starts normally when cold.
  • After about one minute:
  • Only F30005 appears.

This indicates a protection calculation process.

The possible sequence is:

Power ON

↓

Power unit initialization

↓

Current feedback activated

↓

Abnormal current offset detected

↓

Software calculates excessive thermal stress

↓

I²T value increases

↓

F30005 occurs

This behavior strongly suggests:

incorrect current feedback rather than a real overload condition.


5. The Critical Diagnostic Data: W Phase Offset -78.35A

The most important diagnostic information in this case is:

Phase current offset:

U phase:
-0.04 A

V phase:
-0.27 A

W phase:
-78.35 A

The Motor Module rating:

Output current: 18A

But the measured W-phase offset:

-78.35A

is more than four times the rated output current.

This is absolutely abnormal.

A healthy current measurement system normally has:

  • Offset close to 0A
  • Small differences between phases
  • Usually within a fraction of an ampere

A value of -78A means:

The drive believes that W-phase current exists even when the motor is not running.


6. Fault Location Analysis

Based on the diagnostic results:

ComponentEvaluation
U-phase current measurementNormal
V-phase current measurementNormal
W-phase current measurementAbnormal
MotorLower probability
DC-link capacitorPossible but not primary
IGBTAlready replaced
Software parameterLow probability

The fault area is concentrated in the W-phase current feedback path:

W-phase CT sensor

↓

CT power supply

↓

CT output signal

↓

Filtering circuit

↓

Amplifier circuit

↓

ADC input

↓

Control electronics

7. Why Did the Overheating Event Damage the Current Measurement Circuit?

The original failure was caused by:

Cabinet cooling fan failure and excessive internal temperature.

Many repairs focus only on replacing:

  • IGBT
  • Fuse

However, overheating affects many other components.


7.1 Current Sensor Damage

The CT sensor may contain:

  • Hall sensor element
  • Signal conditioning circuit
  • Temperature compensation components

High temperature can cause:

  • Zero-point drift
  • Sensitivity change
  • Output instability

7.2 Analog Circuit Damage

Behind the CT sensor there are usually:

  • Filtering resistors
  • Capacitors
  • Operational amplifiers
  • Protection components

High temperature can cause:

  • Resistor value drift
  • Capacitor leakage
  • Amplifier input damage

7.3 Secondary Damage from IGBT Failure

When an IGBT fails:

The fault current path can be:

IGBT failure

↓

DC bus current surge

↓

Current sensor

↓

Measurement circuit

Even if the IGBT is replaced successfully, the current feedback circuit may remain damaged.


8. Why Replacing the CT Sensor May Not Solve the Problem

Replacing the CT sensor does not guarantee repair.

The following points must be confirmed:

8.1 Correct CT Model

The replacement CT must have:

  • Same model number
  • Same sensitivity
  • Same output characteristics
  • Same temperature compensation

A physically identical sensor may still be electrically different.


8.2 Correct Installation Direction

Hall current sensors are directional.

Incorrect installation can cause:

  • Negative output
  • Incorrect polarity
  • Large current offset

8.3 Correct Wiring

The following must be verified:

  • Positive supply
  • Negative supply
  • Signal output
  • Ground connection

8.4 Calibration Requirements

After replacing power components, some systems may require:

  • Current offset calibration
  • Drive identification procedure

Otherwise, the current measurement may remain incorrect.


9. Recommended Troubleshooting Procedure

Step 1: Disconnect Motor Cable

Remove:

U
V
W

from the motor.

Purpose:

Eliminate:

  • Motor insulation problems
  • Cable short circuit

Step 2: Check Current Offset Parameters

Monitor:

r0069[3]  Phase U offset

r0069[4]  Phase V offset

r0069[5]  Phase W offset

The three values should be close to each other.

A difference of tens of amperes indicates a measurement circuit failure.


Step 3: Measure CT Supply Voltage

Compare:

  • U-phase CT
  • V-phase CT
  • W-phase CT

Measure:

  • Supply voltage
  • Ground reference
  • Output voltage

All three channels should be similar.


Step 4: Check CT Output Signal

With zero current:

The CT output should remain stable.

If W-phase output is:

  • 0V
  • 5V
  • unstable voltage

the sensor or signal circuit is faulty.


Step 5: Inspect PCB Components

Focus on the W-phase measurement area:

  • Solder joints
  • Signal resistors
  • Filter capacitors
  • Operational amplifier
  • PCB traces

High-current IGBT failures often leave hidden damage.


10. Common Repair Mistakes

Mistake 1: Only Replacing IGBT

Many technicians see F30021 and immediately replace IGBT.

However:

The IGBT may only be the damaged component, not the root cause.


Mistake 2: Ignoring Current Feedback

Modern drives depend heavily on feedback signals.

Incorrect feedback can create:

  • False overcurrent
  • False ground fault
  • False thermal overload

Mistake 3: Not Checking Diagnostic Parameters

SINAMICS S120 provides detailed diagnostic information:

Examples:

  • r0069 current feedback
  • r0949 fault values
  • Fault history

Ignoring these parameters makes troubleshooting much more difficult.


11. Final Diagnosis of This Case

Considering all information:

  • Cabinet cooling failure caused overheating.
  • Initial faults were F30021 and F30005.
  • IGBT was replaced.
  • CT sensors were replaced.
  • Fault still appears after approximately one minute.
  • W-phase current offset is -78.35A.

The most likely causes are:

First possibility: W-phase current sensing failure

Probability: approximately 50%

Possible reasons:

  • Incorrect CT installation
  • Wrong CT replacement model
  • Damaged W-phase CT
  • CT power supply failure

Second possibility: W-phase analog feedback circuit damage

Probability: approximately 35%

Possible components:

  • Signal resistor
  • Filter capacitor
  • Operational amplifier
  • ADC input circuit

Third possibility: IGBT driver circuit problem

Probability: approximately 15%


12. Conclusion

SINAMICS S120 F30005 and F30021 faults should not be diagnosed only by replacing power semiconductors.

In high-power industrial drives, the real failure may exist in:

  • Current sensing circuits
  • Gate driver circuits
  • Protection feedback systems

In this case, the most valuable diagnostic information is:

Phase W offset = -78.35A

This proves that the drive detects a huge W-phase current even without normal motor operation.

The correct repair approach is not simply:

Replace IGBT again.

Instead, the troubleshooting path should follow:

Power semiconductor

↓

Gate driver

↓

Current sensor

↓

Signal conditioning circuit

↓

Control feedback

By analyzing the current feedback system, technicians can accurately locate the fault, prevent repeated IGBT failures, and significantly improve the repair success rate of Siemens SINAMICS S120 Motor Modules.

Posted on

Systematic Diagnosis of High-Temperature Shutdowns and Power-Version Read Failures in Antminer S19 XP Hyd. and S19 XP+ Hyd. Miners

Introduction

In large-scale hydro-cooled mining facilities, Antminer S19 XP Hyd. and S19 XP+ Hyd. miners are commonly installed in groups and operated through centralized power distribution and cooling systems. Compared with conventional air-cooled miners, hydro-cooled models eliminate high-speed cooling fans and instead rely on circulating coolant, internal cold plates, pumps, heat exchangers, manifolds, and external cooling equipment to remove heat generated by the ASIC chips.

This architecture offers several advantages, including lower acoustic noise, improved power density, and better thermal performance in demanding environments. However, it also introduces a more complex fault chain. A hydro-cooled miner does not judge its thermal condition simply from room temperature. Its control system continuously evaluates hashboard temperatures, PIC controller temperatures, power-supply communication, ASIC detection results, voltage regulation, frequency ramping, and several protection thresholds.

As a result, maintenance personnel can easily misdiagnose a machine when the miner log simultaneously displays messages such as:

get power version failed
check power version failed, use v2 protocol to try it again
ERROR_TEMP_TOO_HIGH
over max temp
power off hashboard

Some technicians immediately conclude that the power supply is defective after seeing “get power version failed.” Others focus only on the temperature alarm and check the ambient temperature. When the room temperature is found to be only 32–33°C, they may assume that the miner’s temperature-monitoring circuit is faulty.

In reality, these messages may appear during the same startup process but have different levels of importance. They may not have the same root cause, and they should not be assigned the same diagnostic priority.

The correct approach is to analyze the complete startup sequence of the miner and separate normal startup retries, compatibility warnings, recoverable communication events, and the final fault that actually caused the shutdown.

This article presents a systematic diagnostic method for Antminer S19 XP Hyd. and S19 XP+ Hyd. miners experiencing high-temperature shutdowns, power-version reading failures, EEPROM-version warnings, unsuccessful initialization, and interruptions during frequency ramping. The discussion is based on a practical multi-miner case involving both operating and stopped hydro-cooled units.


Technician inspecting Antminer S19 XP Hyd. miners with a tablet displaying high-temperature and miner log warnings in a hydro-cooled mining facility.

1. Basic System Architecture of a Hydro-Cooled Miner

A complete Antminer S19 XP Hyd. or S19 XP+ Hyd. system normally includes the following components:

  1. Control board
  2. Three hydro-cooled hashboards
  3. Dedicated miner power supply
  4. High-current copper busbars or equivalent power connections
  5. Power-supply communication cable
  6. Hashboard signal cables
  7. Coolant distribution manifold
  8. Internal cold plates and coolant channels
  9. Coolant inlet and return hoses
  10. External circulation pump
  11. Heat exchanger or cooling tower
  12. Flow, pressure, and temperature-monitoring devices

The control board performs several critical tasks. It boots the embedded Linux operating system, loads miner firmware, reads hashboard EEPROM data, identifies ASIC chips, communicates with the power supply, controls operating voltage and frequency, receives temperature data, and executes shutdown protection when abnormal conditions are detected.

The hashboards contain large numbers of ASIC chips. The exact ASIC count and operating frequency differ between S19 XP Hyd. and S19 XP+ Hyd. versions, but the overall control logic is similar: multiple chains operate together, and the miner must correctly detect and initialize each chain before normal hashing can begin.

The power supply is not merely a fixed-output DC source. The control board communicates with it digitally to obtain information such as:

  • Power-supply version
  • Serial number
  • Calibration status
  • Calibration data
  • Output-voltage capability
  • Watchdog status
  • Current operating state

The control board also commands the power supply to raise or lower its output according to the miner’s operating frequency and workload.

The hydro-cooling system carries heat away from the ASIC chips and transfers it to the circulating coolant. The heat is then removed by an external radiator, heat exchanger, cooling tower, or other cooling device.

A failure in any one of these subsystems can prevent startup or cause a shutdown during operation.


Close-up of an Antminer S19 XP+ Hyd. miner showing a Chain 2 overtemperature alarm, coolant hoses, and diagnostic log information.

2. Why a Miner Can Report High Temperature When the Room Is Only 32–33°C

One of the most common misunderstandings in hydro-miner troubleshooting is the assumption that a high-temperature alarm refers directly to the room temperature.

For a hydro-cooled miner, room temperature is only one external condition. The protection algorithm primarily reacts to temperatures measured inside the miner.

A room temperature of 32°C does not guarantee that the hashboards are being cooled correctly. Similarly, an inlet coolant temperature of approximately 29–33°C does not prove that coolant is flowing properly through every cold plate.

The heat-transfer path is approximately:

ASIC junction
→ chip package
→ thermal interface material
→ cold plate
→ coolant
→ hose network
→ heat exchanger
→ surrounding environment

Any abnormal thermal resistance along this path can produce a local overtemperature condition.

Typical causes include:

  • Internal blockage in one cold plate
  • Air trapped inside a coolant channel
  • A bent or compressed hose
  • A quick connector that is not fully opened
  • Uneven flow distribution among parallel branches
  • Insufficient pump flow
  • Excessively high coolant viscosity
  • Contaminated coolant
  • Scale or corrosion inside the cold plate
  • Poor contact between the ASIC surface and cold plate
  • Degraded thermal interface material
  • Temperature-sensor drift
  • Faulty PIC temperature acquisition
  • Abnormal local power dissipation

Therefore, normal room temperature does not exclude a genuine internal overtemperature fault.


3. The Correct Way to Read the Miner Log

A miner log should be read as a chronological startup sequence. It should not be interpreted by selecting one alarming line and ignoring the events before and after it.

The startup process can generally be divided into seven stages.


Engineer monitoring multiple hydro-cooled Antminer miners, coolant systems, and mining status from a laptop in an industrial mining facility.

4. Stage One: Control-Board System Startup

During the first stage, the embedded operating system initializes storage, memory, network interfaces, the FPGA, and the web server.

Typical log entries include:

UBIFS mounted
eth0: link becomes ready
start the http server
httpserver:6060 started ok

If the miner’s web interface can be accessed normally, the following components are probably functioning at a basic level:

  • Main processor
  • Flash storage
  • Ethernet interface
  • Operating system
  • Web-management service

This does not prove that the control board is completely fault-free, but it makes a severe CPU, memory, or network failure less likely.

Boot messages such as file-system recovery, reserved block information, or kernel initialization are not necessarily faults. They are often standard system messages.

The important question is whether the system reaches the miner application and begins hardware initialization.


5. Stage Two: Detection of Hashboards and Chains

A normal startup may include:

board num = 3
board id = 0, chain num = 1
board id = 1, chain num = 1
board id = 2, chain num = 1
chain num = 3

This indicates that the control board recognizes three hashboards or three chains.

If only one or two chains are detected, possible causes include:

  • Loose or damaged signal cable
  • Control-board connector failure
  • Missing auxiliary voltage on one hashboard
  • PIC controller not starting
  • Severe short circuit on the hashboard
  • Incompatible firmware
  • Corrupted board data

When all three chains are detected, the basic communication path is present. However, this does not prove that all chips are healthy or that the cooling system is working correctly.


6. Stage Three: Reading Hashboard EEPROM Data

The log may show:

load chain 0 eeprom data
version invalid 5!=4
v5 load data succ

This sequence is often misunderstood.

The message:

version invalid 5!=4

may simply indicate that the firmware first attempted to interpret the EEPROM using one data structure, found a different version number, and then retried with the correct version.

The following line:

v5 load data succ

shows that version-5 data was successfully loaded.

Therefore, the warning by itself does not prove that the EEPROM is damaged.

The message becomes more significant only if it is accompanied by errors such as:

eeprom read failed
crc error
invalid board data
cannot load eeprom
chain data missing

or if the miner fails to determine the correct model, ASIC count, voltage parameters, or target frequency.

In the case discussed here, all three chains reported “version invalid 5!=4” followed by successful version-5 data loading. This suggests a compatibility-handling process rather than a fatal EEPROM failure.


7. Stage Four: ASIC Chip Detection

A healthy S19 XP Hyd. chain may report:

Chain[0]: find 204 asic
Chain[1]: find 204 asic
Chain[2]: find 204 asic

When all three chains detect the full expected number of ASIC chips, several conclusions can be drawn:

  • The control-board-to-hashboard communication is functioning.
  • The clock signal reaches the ASIC chain.
  • The reset path is functioning.
  • Serial communication through the ASIC chain is intact.
  • The PIC controller is operating sufficiently to support initialization.
  • There is no obvious chip-chain break.
  • No major number of chips is missing.

This does not guarantee that the hashboard is completely healthy. A board can still have a cooling fault, unstable voltage rail, marginal ASIC, thermal-sensor fault, or intermittent communication issue.

However, when all 204 ASICs are found on each chain, the probability of a classic “missing chip,” “broken chain,” or signal-line fault becomes much lower.

If a chain detects only part of the expected ASIC count, the technician should instead investigate:

  • Clock signal
  • Reset signal
  • CO, RI, BO, or equivalent chain signals
  • Local voltage domains
  • Damaged ASICs
  • Broken solder joints
  • Shorted or open components
  • Fault location within the chain

8. Stage Five: Communication With the Power Supply

The following sequence may appear:

get power version failed
check power version failed, use v2 protocol to try it again
power open power_version = 0x65
power is calibrated
power sn: XXXXX
enable_power_calibration

This must be interpreted as a complete sequence.

The first line:

get power version failed

means that the first attempt to read the power-supply version did not succeed.

The next line:

use v2 protocol to try it again

shows that the firmware changed its communication method and retried.

If the later log displays:

power open power_version = 0x65
power is calibrated
power sn: XXXXX

then the control board eventually obtained useful data from the power supply, including:

  • Power type or version
  • Serial number
  • Calibration status
  • Calibration record

In this situation, it is incorrect to declare that the power supply is completely dead.

More likely explanations include:

  • The first handshake timed out.
  • The firmware and PSU use different protocol revisions.
  • The PSU started responding later than expected.
  • The communication line has marginal contact.
  • There is electrical interference.
  • The PSU controller has a delayed initialization sequence.
  • The firmware includes a compatibility fallback mechanism.

The severity of “get power version failed” depends on what happens next.

When the miner can later set voltage, detect ASICs, and ramp frequency, the power supply is clearly providing functional output. The message may still deserve attention, but it is not necessarily the direct cause of the shutdown.


9. When “Get Power Version Failed” Is a Serious Fault

The message becomes more important when the following conditions are also present:

get power version failed
power communication timeout
cannot read power SN
power init failed

and when:

  • Hashboards do not receive operating voltage.
  • ASIC detection never starts.
  • Frequency ramping does not begin.
  • The miner repeatedly restarts.
  • The PSU does not respond to voltage commands.
  • The PSU communication data remains completely unavailable.

In that case, the technician should inspect:

  • PSU communication cable
  • Connector pin contact
  • Cable continuity
  • Auxiliary supply voltage
  • Control-board communication port
  • PSU internal controller
  • Firmware compatibility
  • PSU model compatibility
  • Grounding and electrical noise

However, that was not the primary pattern in the high-temperature case. The affected miners were still able to recognize all ASICs and begin frequency ramping.


10. Stage Six: Voltage Establishment and Frequency Ramping

A normal miner does not immediately start at full operating frequency. It usually increases voltage and frequency in controlled steps.

Typical entries include:

set_voltage_by_steps to 2100
Chain[0]: find 204 asic
Chain[1]: find 204 asic
Chain[2]: find 204 asic
chain 0 set freq to 50
chain 1 set freq to 50
chain 2 set freq to 50
...
chain 0 set freq to 475
chain 1 set freq to 475
chain 2 set freq to 475

The gradual ramp serves several purposes:

  • Reduces inrush stress
  • Tests ASIC stability at increasing frequency
  • Confirms adequate operating voltage
  • Monitors chip response
  • Detects thermal abnormalities
  • Allows the firmware to adjust voltage margins
  • Prevents sudden excessive current demand

If the miner successfully ramps from approximately 50 MHz to 475 MHz, several subsystems are working:

  1. The PSU is supplying output power.
  2. The control board can command voltage changes.
  3. The ASIC chains respond to frequency commands.
  4. The hashboards are capable of operating at least during initialization.
  5. The signal paths are functioning.

This is strong evidence against a complete PSU failure.

In the analyzed case, the miner reached high frequency before the temperature protection occurred. Therefore, the final shutdown was not caused by an inability to power the hashboards.


11. Stage Seven: Final Protection Trigger

The decisive log lines were:

over max temp, pic temp(chain2) 81 (max 80)
ERROR_TEMP_TOO_HIGH
stop mining: over max temp
power off hashboard

This is the actual shutdown event.

The log provides four important facts:

  1. The problem occurred on Chain 2.
  2. The measured PIC-related temperature reached 81°C.
  3. The configured maximum was 80°C.
  4. The firmware intentionally shut down the hashboards.

Therefore, the direct reason for stopping was overtemperature protection.

The earlier power-version message was not the final shutdown command.

This distinction is essential. A miner log may contain multiple warnings, but only one event may actually cause the stop command.


12. Case Comparison Between Running and Stopped Miners

The site included multiple Antminer S19 XP Hyd. and S19 XP+ Hyd. miners.

The monitoring interface showed a mixed operating condition:

  • Some miners were producing approximately 259 TH/s.
  • Some were producing approximately 298 TH/s.
  • Some were producing approximately 303 TH/s.
  • Several miners were stopped at 0 GH/s.
  • One stopped miner showed a temperature range near 48–96°C.
  • Another stopped miner showed 0–32°C.
  • Different machines used different firmware compilation dates.
  • Several machines reported similar EEPROM-version messages.
  • Some machines reported initial PSU-version read failures.

This is important because it proves that the entire cooling and power-distribution system was not completely nonfunctional. Several miners were still hashing.

At the same time, it does not prove that every cooling branch was healthy. In a parallel hydro-cooling system, some machines can receive adequate coolant flow while others receive insufficient flow.

The correct diagnostic approach is therefore machine-specific. Each IP address must be analyzed separately.

It is a mistake to assume that all miners of the same model have the same fault merely because they are connected to the same cooling rack.


13. The Significance of the 81°C Chain-2 Alarm

The critical line was:

pic temp(chain2) 81 (max 80)

This does not mean that the room was 81°C. It does not necessarily mean that the bulk coolant reached 81°C. It means that the temperature value associated with Chain 2 and monitored by the firmware crossed the protection threshold.

The monitoring software also showed one miner with a displayed temperature range up to 96°C. This is consistent with a local thermal problem or a temperature-sensing abnormality.

Because the fault repeatedly involved a specific chain rather than all three chains simultaneously, the most likely problem is local rather than system-wide.

Potential local causes include:

  • Restricted flow through the Chain-2 cold plate
  • Air pocket in the Chain-2 coolant channel
  • Partially closed quick connector
  • Collapsed hose
  • Internal cold-plate blockage
  • Poor thermal contact
  • Faulty temperature sensor
  • PIC acquisition error

If all three chains had shown a similar high temperature at the same time, a general cooling-system fault would be more likely.


14. How to Distinguish Real Overtemperature From False Temperature Detection

Before replacing hardware, the technician must determine whether the reported temperature is physically real.

14.1 Observe the Temperature Rise Rate

A very rapid increase from approximately 30°C to 80°C within seconds or a few minutes suggests:

  • No coolant flow
  • Severe blockage
  • Closed valve
  • Quick connector not open
  • Air lock
  • Faulty sensor

A slower temperature increase over a longer period suggests:

  • Insufficient total pump flow
  • High coolant inlet temperature
  • Heat exchanger capacity too low
  • Cooling-tower performance problem
  • Poor flow balancing among multiple miners
  • Excessive total heat load

The rate of temperature rise is often more informative than the final temperature alone.


14.2 Compare All Three Chains

The temperatures of the three hashboards should be reasonably similar under equal operating conditions.

For example:

Chain 0: 55°C
Chain 1: 57°C
Chain 2: 81°C

This pattern strongly indicates a Chain-2-specific issue.

By contrast:

Chain 0: 79°C
Chain 1: 80°C
Chain 2: 82°C

would suggest a broader cooling-capacity or inlet-temperature problem.

A single-chain deviation is usually caused by local flow restriction, local heat-transfer resistance, or sensor error.


14.3 Measure Inlet and Outlet Temperatures

The following temperatures should be measured:

  • Main system supply temperature
  • Main system return temperature
  • Individual miner inlet temperature
  • Individual miner outlet temperature
  • Individual branch temperatures where practical

An infrared thermometer may be used, but transparent or translucent hoses can produce inaccurate readings. A better method is to attach a piece of matte black tape to the hose and measure the tape surface.

A contact temperature probe or thermal camera is preferable when available.

Interpretation examples:

High board temperature with almost no inlet-to-outlet temperature difference

Possible explanations:

  • Coolant is not passing through the cold plate.
  • A bypass path exists.
  • A quick connector is closed.
  • The internal channel is blocked.
  • The reported temperature is false.

Excessively large inlet-to-outlet temperature difference

Possible explanations:

  • Coolant flow is too low.
  • Hydraulic resistance is too high.
  • A hose or cold plate is partially blocked.
  • Pump pressure is inadequate.

High inlet temperature on all miners

Possible explanations:

  • Heat exchanger insufficient
  • Cooling tower failure
  • High ambient wet-bulb temperature
  • Pump circulation problem
  • Total system heat load above design capacity

15. Flow Measurement Is More Reliable Than Visual Inspection

The coolant hoses in the site photos appeared connected, but visual inspection alone cannot confirm flow.

A hose can look normal while containing almost no coolant movement. Similarly, a quick connector can appear fully inserted while its internal valve remains partially closed.

The best verification methods are:

  • Inline flow meter
  • Differential pressure measurement
  • Clamp-on ultrasonic flow meter
  • Comparison of branch pressure
  • Controlled return-flow test
  • Temporary transparent inspection section

The technician should compare the faulty branch with a known-good branch.

A practical test is to record:

  • Flow rate of a normal miner
  • Flow rate of the stopped miner
  • Inlet temperature
  • Outlet temperature
  • Time to overtemperature trip

The comparison can quickly reveal whether the fault is hydraulic.


16. Cold-Plate Internal Blockage

Hydro-cooled mining systems can develop internal contamination over time.

Possible contaminants include:

  • Metal oxides
  • Corrosion products
  • Scale
  • Sealant debris
  • Hose particles
  • Biological growth
  • Improperly mixed coolant residue
  • Foreign matter introduced during maintenance

Cold plates often contain narrow internal passages. Even partial restriction can reduce flow enough to cause local overheating at full frequency.

The miner may appear normal at low frequency because heat generation is limited. During frequency ramping, thermal load increases rapidly. Once the miner reaches approximately 400–475 MHz, the restricted cooling path can no longer remove the generated heat, and the temperature rises above the threshold.

This pattern matches a miner that passes ASIC detection and low-frequency initialization but trips near the end of the ramp.

A blocked cold plate should be compared hydraulically with a healthy one. Any cleaning procedure must use a compatible fluid and controlled pressure. Excessive pressure can damage seals or deform the cold plate.


17. Air Lock and Trapped Gas

Air inside the coolant loop is a frequent cause of local overheating.

Air can accumulate at high points, inside the cold plate, or near a connector. Because air has much lower heat-transfer capability than liquid coolant, the affected area can overheat even while the main pump is running.

Symptoms include:

  • Visible bubbles
  • Irregular coolant noise
  • Sudden temperature fluctuations
  • One board heating much faster than the others
  • Temperature falling rapidly after shutdown
  • Temporary improvement after hose movement
  • Unstable flow-meter readings

Corrective actions include:

  • Proper system bleeding
  • Opening high-point vents
  • Checking expansion-tank level
  • Reorienting hoses to avoid air pockets
  • Running the pump at controlled speed during bleeding
  • Ensuring the pump does not draw air
  • Verifying adequate reservoir level

A hydro-mining system should not be restarted at full load until trapped air has been removed.


18. Quick Connector Not Fully Open

Many liquid-cooling systems use self-sealing quick connectors. These connectors contain internal valves that open only when the connector is fully seated.

A connector may appear connected while the internal valve remains only partially open.

In that condition:

  • Some coolant may still pass.
  • The hose may feel cool.
  • No leak may be visible.
  • Low-frequency operation may appear normal.
  • Full-load operation may trigger overtemperature.

The connector should be disconnected only after the system is powered down and depressurized. It should then be inspected for:

  • Incomplete engagement
  • Damaged locking mechanism
  • Bent internal valve
  • Seal swelling
  • Internal contamination
  • Flow restriction

Never disconnect a pressurized coolant line while the miner is energized.


19. Hose Kinking and Mechanical Restriction

The site photographs showed a dense arrangement of multiple hoses around stacked miners. In such installations, hose routing is critical.

Common mechanical restrictions include:

  • Excessively tight bend radius
  • Hose trapped under a metal frame
  • Hose compressed by another hose
  • Soft hose collapsing at elevated temperature
  • Twisted reinforcement layer
  • Debris inside the hose
  • Long unsupported hose creating a low point

The full hose length must be inspected, not only the connector ends.

A hose that is only partially flattened can still pass coolant at low demand but become inadequate during full-load operation.


20. Uneven Flow Distribution in Parallel Systems

When multiple hydro miners are connected in parallel, coolant does not automatically divide equally among all branches.

Fluid follows the paths of lowest hydraulic resistance. A miner close to the pump or manifold may receive more flow, while a distant branch or a branch with additional restrictions may receive less.

This can produce a mixed condition:

  • Several miners run normally.
  • Several miners repeatedly overheat.
  • The same rack positions are affected.
  • Increasing pump speed improves the fault.
  • Stopping some miners allows others to recover.

Correct system design may require:

  • Balancing valves
  • Individual flow meters
  • Equal-length piping
  • Reverse-return layout
  • Properly sized manifolds
  • Adequate pump head
  • Pressure monitoring at supply and return

A centralized system should not be judged only by total flow. The branch flow to each miner is the important value.


21. Poor Thermal Contact Between ASICs and Cold Plate

Adequate coolant flow does not guarantee adequate heat transfer from the ASICs.

The thermal path between the ASIC package and the cold plate may be compromised by:

  • Aged thermal pads
  • Incorrect thermal-pad thickness
  • Uneven cold-plate pressure
  • Loose fasteners
  • Warped cold plate
  • Warped PCB
  • Contamination between surfaces
  • Damaged thermal interface material
  • Previous improper repair

In such a case, the coolant may remain relatively cool because heat is not transferring effectively into it, while the ASIC or board temperature rises rapidly.

This type of fault is more difficult to confirm externally. If flow is verified and the sensor reading is believable, the hashboard may require removal and inspection by a qualified repair technician.

The cold plate should not be opened casually because improper reassembly can cause coolant leakage directly onto energized electronics.


22. Temperature Sensor or PIC Acquisition Fault

The log refers specifically to a PIC-related temperature value on Chain 2.

If physical temperature measurements show that the board and coolant are normal, but the firmware repeatedly reports 81–96°C, a sensor or acquisition fault should be considered.

Possible causes include:

  • Temperature sensor drift
  • Open sensor circuit
  • Shorted sensor circuit
  • Incorrect pull-up or bias resistor
  • ADC-reference abnormality
  • PIC firmware error
  • Local auxiliary-supply problem
  • Damaged signal trace
  • Moisture or corrosion

A sensor fault may produce:

  • Unrealistically fixed temperature
  • Sudden jumps
  • Large difference from external measurement
  • Same value after a cold restart
  • Immediate overtemperature before power ramping
  • Temperature reading inconsistent with coolant temperature

The technician can compare the suspect board with a normal board by:

  • Swapping the control board
  • Swapping hashboard signal cables
  • Moving the suspect hashboard to another compatible unit
  • Reading sensor voltages
  • Comparing PIC data
  • Using service software if available

Only one variable should be changed at a time.


23. The Importance of Cross-Swap Testing

Cross-swap testing is one of the most effective diagnostic tools in a multi-miner facility.

Possible swap tests include:

  • Faulty miner cooling branch with a known-good branch
  • PSU communication cable
  • Control board
  • Power supply
  • Hashboard signal cable
  • Entire hashboard, when technically feasible

The interpretation is straightforward.

Fault follows the cooling branch

Likely causes:

  • Hose restriction
  • Quick connector problem
  • Manifold imbalance
  • Valve problem
  • Inadequate local flow

Fault stays with the same hashboard

Likely causes:

  • Cold-plate blockage
  • Poor thermal contact
  • Temperature sensor fault
  • PIC acquisition fault
  • Local electrical problem

Fault follows the control board

Likely causes:

  • Firmware problem
  • Temperature-data interpretation fault
  • Control-board communication fault
  • Corrupted configuration

Fault follows the power supply

Likely causes:

  • PSU communication instability
  • Incorrect output regulation
  • PSU controller issue
  • Calibration-data problem

Every swap must be documented. Changing multiple components simultaneously destroys the diagnostic value of the test.


24. Firmware Inconsistency Across the Site

The supplied logs showed different firmware compilation dates and versions.

Some miners used firmware compiled in 2024, while others used builds from 2025. One S19 XP+ Hyd. miner displayed a later FR-series firmware version.

This matters because firmware controls:

  • Power-supply protocol
  • EEPROM interpretation
  • ASIC count and chain parameters
  • Voltage tables
  • Frequency tables
  • Temperature thresholds
  • Calibration procedures
  • Startup timing
  • Automatic recovery behavior

S19 XP Hyd. and S19 XP+ Hyd. are similar but not interchangeable at the firmware level.

Using the wrong firmware can cause:

  • PSU-version communication warnings
  • EEPROM-format mismatch
  • Incorrect working voltage
  • Incorrect target frequency
  • Incorrect temperature threshold
  • Missing calibration files
  • Initialization loops
  • Automatic restart
  • Reduced hashrate
  • Unstable operation

Firmware should only be changed after verifying:

  • Exact miner model on the nameplate
  • Rated hashrate
  • Control-board hardware version
  • Power-supply model
  • Current firmware version
  • Official firmware applicability

A firmware update should not be the first response to an obvious single-chain cooling fault. Unnecessary flashing may introduce additional variables and make diagnosis more difficult.


25. Interpreting “Version Invalid 5!=4”

The repeated log entry:

version invalid 5!=4
v5 load data succ

is not, by itself, a reason to replace the hashboard or rewrite the EEPROM.

The successful follow-up line is the key.

The miner subsequently:

  • Identified the correct machine type
  • Recognized all three chains
  • Detected all ASIC chips
  • Created mining threads
  • Set voltage
  • Ramped frequency

That means the EEPROM data was usable.

The warning should be treated as a compatibility message unless the miner also shows incorrect board parameters or failure to initialize.


26. A Separate Initialization Fault on Another Miner

One of the S19 XP+ Hyd. logs showed a different pattern:

fail to open counter file, err: No such file or directory
ERROR_INIT: something error when miner init, restart...
stop_mining_and_restart
power off hashboard
restart

This is not the same as the Chain-2 high-temperature event.

The miner recognized the boards and read PSU calibration data, but then failed to open a required file and restarted.

Possible causes include:

  • Missing firmware file
  • Damaged configuration partition
  • Incomplete firmware installation
  • Corrupted calibration-related data
  • File-system problem
  • Incompatible firmware build
  • Control-board flash-memory issue

This machine should be handled separately from the high-temperature units.

A suitable process would be:

  1. Record current settings.
  2. Download the complete log.
  3. Verify the exact machine model.
  4. Restore factory settings.
  5. Install the correct official firmware.
  6. Perform a full cold restart.
  7. Observe whether the missing-file error returns.
  8. Cross-test the control board if necessary.

This demonstrates why each IP address requires its own fault classification.


27. Recommended Standard Diagnostic Procedure

Step 1: Record the Machine Identity

For every miner, record:

  • IP address
  • Model
  • Serial number
  • Rated hashrate
  • Firmware version
  • Current hashrate
  • Chain temperatures
  • PSU model
  • Last shutdown time
  • Final 100–200 log lines

This is especially important when several units fail at the same site.


Step 2: Identify the Final Shutdown Cause

Search the end of the log for messages such as:

ERROR_TEMP_TOO_HIGH
ERROR_POWER
ERROR_INIT
ERROR_SOC_INIT
ERROR_ASIC
stop mining
power off hashboard
restart

The final protection event normally has higher diagnostic priority than an earlier warning.


Step 3: Verify Board and ASIC Detection

Confirm:

  • Three chains detected
  • Expected ASIC count found on each chain
  • No repeated chain-break errors
  • No missing-board condition

When all ASICs are found, do not immediately assume the hashboard electronics are defective.


Step 4: Inspect the Hydro-Cooling System

Check:

  • Pump running status
  • Total system flow
  • Branch flow
  • Supply temperature
  • Return temperature
  • System pressure
  • Reservoir level
  • Coolant condition
  • Air bubbles
  • Quick-connector engagement
  • Hose kinks
  • Valve position
  • Manifold balance

Step 5: Compare With a Working Miner

A nearby working miner is an excellent reference.

Compare:

  • Firmware version
  • PSU version
  • Startup sequence
  • ASIC count
  • Frequency-ramp time
  • Inlet and outlet temperatures
  • Branch flow
  • Operating temperature
  • PSU communication messages

A good reference unit can reveal whether a log message is normal for that firmware or specific to the failed machine.


Step 6: Perform Controlled Cross-Swap Tests

Swap only one item at a time and record the result.

Recommended order:

  1. Cooling branch
  2. Hose and quick connector
  3. Signal cable
  4. PSU communication cable
  5. Control board
  6. Power supply
  7. Hashboard

The order may be adjusted according to site conditions, but the principle remains the same.


Step 7: Address Firmware Only After Hardware Checks

Firmware restoration or upgrading should be considered when:

  • Water flow is verified.
  • Physical temperatures are normal.
  • Connections are secure.
  • The fault remains repeatable.
  • Initialization or file errors appear.
  • PSU protocol mismatch persists.
  • Control-board cross-test implicates firmware or storage.

28. Safety Precautions

Hydro-cooled miners operate at high power and high current. Maintenance work involves both electrical and liquid-cooling hazards.

Essential precautions include:

  1. Switch off the miner completely before opening the unit.
  2. Wait for internal capacitors to discharge.
  3. Do not connect or disconnect hashboard cables while energized.
  4. Do not disconnect coolant hoses while pressurized.
  5. Prevent coolant from entering the PSU or control board.
  6. Use a collection tray before opening the cooling loop.
  7. Verify protective grounding.
  8. Tighten high-current terminals to the correct torque.
  9. Do not repeatedly power-cycle the miner at short intervals.
  10. Never bypass high-temperature protection.
  11. Do not force full-frequency operation before confirming coolant flow.
  12. Wear appropriate electrical and eye protection.

High-temperature shutdown is a protective response, not a nuisance alarm. Defeating it may cause permanent ASIC damage, cold-plate deformation, solder-joint failure, or fire risk.


29. Most Probable Root Causes in This Case

Based on the complete logs and operating sequence, the likely causes can be prioritized as follows.

First Priority: Local Cooling-Flow Restriction

Examples:

  • Partially blocked hose
  • Quick connector not fully open
  • Internal cold-plate restriction
  • Trapped air
  • Uneven manifold flow
  • Closed or partially closed valve

This is the most likely category because the fault affected a specific chain and occurred during high-frequency ramping.


Second Priority: Temperature-Sensing Error

Examples:

  • Sensor drift
  • PIC acquisition error
  • ADC-reference problem
  • Sensor wiring issue
  • Local auxiliary-voltage instability

This becomes more likely if external temperature and flow measurements do not support the 81–96°C reading.


Third Priority: Poor Thermal Interface

Examples:

  • Degraded thermal pad
  • Uneven cold-plate pressure
  • Loose mounting
  • Board or plate deformation
  • Contaminated interface

This can cause a high local temperature even when coolant flow is normal.


Fourth Priority: Firmware or Compatibility Issue

Firmware differences may contribute to:

  • Power-version read retries
  • EEPROM-version messages
  • Different temperature-handling logic
  • Calibration-file errors

However, firmware is less likely to be the primary cause of a repeatable single-chain physical overtemperature event.


Fifth Priority: Complete PSU Failure

A complete PSU failure is unlikely in a miner that:

  • Reads PSU calibration data
  • Finds all ASICs
  • Raises voltage
  • Ramps frequency to 475 MHz
  • Stops only after an overtemperature command

The PSU communication warning should still be monitored, but it is not the main shutdown cause in this sequence.


30. Validation After Repair

A repaired miner should not be considered healthy merely because it begins hashing.

The following items should be verified:

  1. All three chains are online.
  2. The full ASIC count is detected.
  3. Chain temperatures remain reasonably balanced.
  4. No chain approaches the thermal limit during startup.
  5. Hashrate reaches the expected range.
  6. The miner runs for at least 30 minutes without shutdown.
  7. The miner runs for at least two hours without continuous temperature rise.
  8. The log contains no repeated PSU communication timeouts.
  9. Inlet and outlet temperatures remain stable.
  10. Coolant flow remains stable.
  11. Pump current and system pressure remain normal.
  12. No leakage occurs at connectors or cold plates.

A miner that operates for only a few minutes after repair has not been fully validated.

If it overheats again after a longer period, the issue may involve inadequate total cooling capacity rather than a local blockage.


31. Practical Diagnostic Logic

A useful diagnostic decision sequence is:

Control board boots?
→ Yes

Three chains detected?
→ Yes

Full ASIC count detected?
→ Yes

Power supply eventually communicates?
→ Yes

Voltage established?
→ Yes

Frequency ramps?
→ Yes

Final error is overtemperature?
→ Yes

Focus on local coolant flow, cold plate, thermal contact, or sensor

By contrast:

Control board boots?
→ Yes

Power version never read?
→ Yes

No PSU serial number?
→ Yes

No voltage established?
→ Yes

No ASIC initialization?
→ Yes

Focus on PSU communication, PSU controller, cable, or firmware compatibility

These two fault paths should not be mixed.


Conclusion

Diagnosing Antminer S19 XP Hyd. and S19 XP+ Hyd. miners requires more than reading a single error message.

The message “get power version failed” does not automatically mean the PSU is defective. The message “version invalid 5!=4” does not automatically mean the hashboard EEPROM is corrupted. A room temperature of 32°C does not prove that the internal cooling system is functioning correctly.

The most reliable approach is to analyze the full startup chain:

Control-board boot
→ hashboard detection
→ EEPROM loading
→ ASIC detection
→ PSU communication
→ voltage establishment
→ frequency ramping
→ temperature monitoring
→ protection shutdown

In the analyzed case, the miners detected all three hashboards, found the full ASIC count, obtained PSU calibration information, established operating voltage, and ramped frequency toward the target. The final event was:

pic temp(chain2) 81 (max 80)
ERROR_TEMP_TOO_HIGH
power off hashboard

This clearly indicates that the immediate cause of shutdown was Chain-2 overtemperature protection.

The troubleshooting priority should therefore move away from immediate PSU replacement and toward:

  • Chain-2 coolant flow
  • Hose condition
  • Quick connectors
  • Air removal
  • Cold-plate restriction
  • Thermal interface quality
  • Temperature-sensor accuracy

For multi-miner hydro-cooled facilities, branch-flow balancing, firmware standardization, accurate maintenance records, and controlled cross-swap testing are essential. By combining log analysis, hydraulic inspection, electrical verification, and comparative testing, technicians can avoid unnecessary power-supply replacement, unnecessary firmware flashing, and premature hashboard disassembly.

A disciplined diagnostic process not only reduces downtime but also prevents secondary damage to expensive ASIC hardware and high-current power components.

Posted on

Diagnostic Methods for Huazhong CNC HNC-210B Spindle Failure, Low-Speed Crawling After Stop, and “Output Shows Low but Relay Remains Energized” — A Practical Analysis Based on PLC/PMC I/O Mapping, Emergency Stop Logic, and the XS91 Spindle Interface

1. Fault Background and Core Problem

The Huazhong CNC HNC-210B is an older integrated CNC control system commonly used on grinders, lathes, special-purpose machines, and retrofit equipment. It combines CNC control, PMC/PLC logic, operator panel keys, machine I/O, spindle interface, feed-axis interface, and related machine control functions in one unit.

In real maintenance work, this type of system often produces faults such as:

The spindle continues to rotate slowly after a stop command.
A relay remains energized, but the CNC diagnostic screen shows the related output point as low level.
The CNC screen shows “Error” or “Drive not ready.”
After entering M03 S100 in MDI, the screen shows M=3 and S=100, but the spindle does not start.
The customer says “one point is faulty” and requests modification of the internal PLC/PMC input/output point assignment.

These faults are easy to misdiagnose. Many technicians may immediately suspect the spindle inverter, relay, CNC output board, or operator panel key. However, the HNC-210B spindle control chain is relatively long. Before judging the hardware, the technician must separate the command layer, PMC signal layer, output point layer, interface voltage layer, external relay layer, and inverter terminal layer.

The key questions are:

Has the CNC accepted the M03 spindle command?
Has the CNC sent a spindle forward request to the PMC/PLC?
Has the PMC/PLC driven Y1.0, Y1.2, or other spindle output points to the XS91 spindle interface?

Only after these layers are separated can the technician determine whether the fault is caused by PLC interlock logic, incorrect I/O mapping, damaged output hardware, external relay feedback, or inverter-side wiring.


HNC-210B CNC controller spindle output test diagram showing XS91 interface wiring, 24V LED test lamps, and multimeter measurement points for spindle forward, enable, and analog speed signals.

2. Basic Structure of the HNC-210B Spindle Control Chain

The spindle on an HNC-210B system is usually not controlled by a direct hardwired output from the spindle key. The signal generally passes through the following chain:

MDI command or spindle key
→ CNC internal spindle command state
→ CNC-to-PMC F signal
→ PMC/PLC logic or standard PLC configuration logic
→ Y output point
→ XS91 spindle interface
→ external relay, inverter start terminal, analog speed reference
→ spindle motor operation

The key signal groups are:

F signals: CNC outputs to PMC/PLC.
G signals: PMC/PLC inputs to CNC.
X signals: machine input points.
Y signals: machine output points.
R, P, B and similar variables: internal relays, parameters, or auxiliary variables.

Common spindle-related signals include:

F06.4: spindle servo enable request.
F06.7: spindle forward request.
F07.0: spindle reverse request.
F07.4: CNC alarm status.
G01.2: spindle lock.
Y1.0: spindle forward output.
Y1.1: spindle reverse output.
Y1.2: spindle enable output.
Y[12] and Y[13]: spindle analog speed reference data.

When the MDI screen shows M=3 and S=100, it only proves that M03 S100 has entered the current CNC state. It does not prove that Y1.0 must be output. To determine whether the CNC actually permits spindle rotation, F06.7 must be checked. To determine whether the PMC is actually driving the hardware terminal, Y1.0 and Y1.2 must be checked. To determine whether the interface hardware is normal, the XS91 connector must be measured directly.


Industrial technician diagnosing a Huazhong HNC-210B CNC controller on a workshop bench while checking PLC and PMC signals, spindle I/O status, and external wiring.

3. Emergency Stop Must Be Cleared Before Spindle Testing

When testing an HNC-210B on the bench, the emergency stop chain must be cleared first. Otherwise, spindle outputs, axis movement, analog output, and some PLC outputs may be inhibited by safety logic. Any test result obtained under an active emergency stop condition is unreliable.

On this machine, the operator panel has an emergency stop button. The rear XT1 terminal also contains external emergency stop contacts. According to the XT1 terminal marking:

1–2: power off.
3–4: power on.
5–6: first emergency stop normally closed contact.
7–8: second emergency stop normally closed contact.

For bench testing, terminals 1–2 and 3–4 usually do not need to be wired if the CNC screen is already powered and running. The important part is the emergency stop chain at terminals 5–6 and 7–8.

When the emergency stop button is released, 5–6 should be closed and 7–8 should also be closed. For bench testing, they can be temporarily shorted as two independent circuits:

XT1-5 to XT1-6
XT1-7 to XT1-8

Do not short terminals 5, 6, 7, and 8 together as one group. They should be treated as two independent normally closed emergency stop contacts.

In addition, the HNC-210B may monitor emergency stop through the handheld unit connector XS8. If no handwheel or handheld unit is installed, the system may remain in emergency stop state. In some machines, the handheld emergency stop input, such as X8.7 to 24VG, must be shorted through a DB25 plug. In this case, the emergency stop was cleared only after shorting XS10 X1.7 to 24VG, indicating that the machine PLC used X1.7 as an external run-permission or emergency stop reset-related input.

These short circuits are for bench diagnosis only. They must not be used as the final safety circuit after the machine is returned to service. The original emergency stop chain must be restored before the equipment is operated.


Engineer repairing an open HNC-210B CNC control unit with a digital multimeter, testing internal circuit boards, XS91 spindle interface signals, and PLC configuration files.

4. Relationship Between X/Y/Z “Drive Not Ready” Alarms and Spindle Testing

After clearing the emergency stop, the CNC may display alarms such as:

25: X-axis drive not ready.
26: Y-axis drive not ready.
27: Z-axis drive not ready.

This is normal on the bench if the X, Y, and Z servo drives are not connected. The feed-axis interfaces such as XS30, XS31, and XS32 receive drive-ready feedback, alarm feedback, encoder signals, or pulse/direction signals. Without the actual drives, the CNC will naturally report that the axis drives are not ready.

However, the spindle can theoretically be tested independently. The XS91 spindle interface and the feed-axis interfaces are separate hardware channels. It is not correct to assume that the spindle hardware cannot be tested simply because the X/Y/Z drives are missing.

The correct question is whether the machine PLC logic uses CNC alarm status F07.4 as a condition to inhibit spindle start.

If F07.4 is 1 and F06.7 is 0, then even if the MDI screen shows M=3, the CNC has not actually sent a spindle forward request to the PMC. In this state, the spindle not starting is not direct proof that the XS91 interface is faulty. It means the normal start path is blocked by alarm state or interlock logic.

Therefore, two different issues must be separated:

Whether the spindle can start under normal machine logic.
Whether the spindle output hardware can be tested independently.

The first depends on alarms, emergency stop, spindle lock, feed drive readiness, machine mode, safety door, hydraulic pressure, lubrication, and other machine conditions. The second can be tested by forced output, LED test lamps, and direct voltage measurement.


5. Key Pins of the XS91 Spindle Interface

The XS91 interface is the key connector for spindle control on the HNC-210B. It is important to confirm that XS91 is a DB15 spindle connector, not a 26-pin connector.

Important XS91 pins are:

Pin 1: AOUT1, -10 V to +10 V analog output.
Pin 2: AGND, analog ground.
Pin 3: +24 V.
Pin 4: Y1.0, spindle forward.
Pin 5: Y1.2, spindle enable.
Pin 8 and pin 10: 24VG.
Pin 9: AOUT2, 0 V to +10 V analog output.
Pin 12: Y1.1, spindle reverse.
Pin 13: Y1.3, spindle orientation.
Pin 14: X3.3, orientation complete input.
Pin 15: X3.2, zero-speed reached input.

The most important signals for basic spindle diagnosis are:

Y1.0: spindle forward output.
Y1.1: spindle reverse output.
Y1.2: spindle enable output.
AOUT1 / AOUT2: spindle speed reference analog output.

If the customer reports that the system shows low level but the relay remains energized, Y1.0 or Y1.2 must be checked. If the customer reports that the spindle continues to rotate slowly after stop, AOUT1 and AOUT2 must be checked after M05 to confirm whether the analog output returns to zero.


6. Correct Bench Wiring for Spindle Output Testing

During bench testing, do not connect the CNC directly to the motor, inverter power terminals, contactor coil, or any unknown field circuit. Use only a 24 V LED test lamp and a multimeter at first.

To test Y1.0 spindle forward output:

XS91 pin 3 +24 V
→ 24 V LED test lamp
→ XS91 pin 4 Y1.0

To test Y1.2 spindle enable output:

XS91 pin 3 +24 V
→ 24 V LED test lamp
→ XS91 pin 5 Y1.2

This wiring is suitable for a common NPN sinking output structure. When the output is active, the Y point sinks to 24VG and the LED turns on. When the output is inactive, the LED remains off.

To test analog output:

Set the multimeter to DC voltage.
Red probe to XS91 pin 1, black probe to XS91 pin 2, to test AOUT1.
Red probe to XS91 pin 9, black probe to XS91 pin 2, to test AOUT2.

Do not randomly short AGND to 24VG or chassis ground. The analog ground and digital ground may be treated differently inside the CNC. Incorrect grounding may cause abnormal analog output or damage the interface.

Minimum test records should include:

XS91 pin 1 to pin 2 voltage during M03 S100.
XS91 pin 9 to pin 2 voltage during M03 S100.
XS91 pin 1 to pin 2 voltage after M05.
XS91 pin 9 to pin 2 voltage after M05.
Whether the Y1.0 LED turns on.
Whether the Y1.2 LED turns on.

Normally, after M03, the spindle forward output should activate and the analog voltage should correspond to the S command. After M05, the spindle forward output should turn off and the analog voltage should return close to 0 V. If the analog output remains several hundred millivolts or more after M05, some inverters may interpret that as a speed command and cause low-speed spindle crawling.


7. Status Diagnosis: Why the Spindle Does Not Start After M03

In this case, after entering M03 S100 in MDI, the screen showed:

M = 3
S = 100

This proves that the command was accepted by the CNC interface. However, the F status showed:

F[006] = 0000H
F[007] = 0F1AH

Bit interpretation:

F06.4 = 0, spindle enable request is not issued.
F06.7 = 0, spindle forward request is not issued.
F07.4 = 1, CNC alarm is active.

This means that although the screen shows M=3, the CNC has not issued the spindle forward request to the PMC. Therefore, the PLC will not drive Y1.0, and XS91 will not output a spindle forward signal.

The G status also showed:

G[000] = 0000H
G[001] = 0000H

This eliminates G01.2 spindle lock as the cause. If G01.2 were 1, the system would lock spindle forward, reverse, and orientation commands. In this case, G01.2 was 0, so the PMC was not actively sending a spindle lock command to the CNC.

Therefore, the logic conclusion is:

M03 has entered the CNC current state.
G01.2 spindle lock is not active.
CNC alarm F07.4 is active.
F06.7 spindle forward request is not issued.
Y1.0 and Y1.2 will not activate.

The normal M03 start path is blocked by system alarm or machine interlock. At this point, it is not correct to immediately condemn the XS91 output hardware. The hardware must be verified separately by forced output or direct I/O mapping tests.


8. Why the Traditional “Ladder Diagram” Interface Cannot Be Found

Many technicians expect to find “Ladder Monitor,” “Ladder Diagnosis,” or “Online Force” directly inside the CNC diagnostic menu. However, on some HNC-210B systems, pressing the Diagnosis key only shows:

Alarm display.
Fault history.
Servo tuning.
Version.
Machining information.
Input/output.
Status display.

Even pressing PgDn does not show ladder monitoring. This does not necessarily mean there is no PLC. It only means the PLC function is not located under the Diagnosis menu.

In this case, the correct path was:

Setting
→ PLC
→ input password HIG
→ PLC file management or Standard PLC Configuration System

After entering this menu, the system displayed disks and PLC files such as:

PLC210B.COM
PLC210B.CPP
PLCTAB_M.DAT
PLC_MAP.H
PMESSAGE.TXT
PARA1.BAP
PARM007.STR

These files prove that PLC logic and configuration data exist in the system.

Their likely meanings are:

PLC210B.COM: compiled PLC runtime file; do not modify directly.
PLC210B.CPP: PLC source file or C-like configuration source.
PLCTAB_M.DAT: PLC point configuration table, likely related to the standard PLC I/O mapping.
PLC_MAP.H: PLC I/O mapping header file.
PMESSAGE.TXT: PLC alarm or message text file.

This type of older system may not provide a Siemens-style or Mitsubishi-style ladder diagram display directly on the CNC screen. It may use a “Standard PLC Configuration System + PLC files + I/O mapping table” structure. The operator can modify standard I/O assignments on the screen, while full source program editing may require Huazhong-specific PLC software, CF card files, or original machine builder project files.


9. External I/O Point Definition in the Standard PLC Configuration System

After entering HIG, the system opens the “Standard PLC Configuration System.” Several configuration pages can be displayed, including:

Spindle system options.
Feed drive options.
Machine frame function options.
Other function options.
Spindle gear shifting.
External I/O point definition.

The “External I/O Point Definition” page is the most important page for this fault. When the customer says “modify the internal point,” this is very likely what they mean. It is not necessarily a traditional ladder program modification.

Typical items in the external I/O definition table include:

External run permission.
SV power ON.
X-axis drive alarm.
Y-axis drive alarm.
Z-axis drive alarm.
X-axis ready.
Y-axis ready.
Z-axis ready.
Spindle forward.
Spindle stop.
Spindle reverse.
Spindle enable.
Spindle orientation.
Spindle zero speed.
Spindle speed reached.
Cooling, lubrication, hydraulic pressure, air pressure, tool magazine, and other machine functions.

Each item usually has:

Input point: group, bit, valid level.
Output point: group, bit, valid level.

The meaning is:

Group: X or Y group number.
Bit: bit number from 0 to 7.
Valid: H or L, meaning high-active or low-active.

Examples:

Output group = 1, bit = 0 means Y1.0.
Output group = 1, bit = 2 means Y1.2.
Input group = 9, bit = 0 means X9.0.
Input group = 1, bit = 7 means X1.7.

If spindle forward was originally assigned to Y1.0 and Y1.0 is confirmed damaged, the spindle forward output can potentially be moved to a spare Y point. The external wiring must then also be moved from the old physical terminal to the new physical terminal. Changing only the software or only the wiring will not fix the machine.


10. Confirm the Output Point Is Really Faulty Before Changing I/O Mapping

Before modifying any I/O point, the output must be confirmed faulty. Otherwise, a normal point may be moved unnecessarily, creating additional faults.

Y output diagnosis should include the following checks:

First, check the Y status display.
For example, Y[001] = 00000000B means Y1.0 to Y1.7 are all 0. If the system shows Y1.0 = 0 but XS91 pin 4 still energizes the relay, possible causes include output transistor leakage, optocoupler damage, external backfeed, or wrong point identification.

Second, test the output with an LED lamp.
Connect the LED test lamp between +24 V and the Y output point. If the Y status is 0, the LED should be off. If the Y status is 1, the LED should be on.

Third, disconnect the external load and test again.
If the output behaves normally with an LED lamp but abnormally when connected to the customer relay, the external circuit may have backfeed, relay coil residual voltage, leakage through a surge suppressor, or voltage coming from the inverter terminal circuit.

Fourth, test analog output residual voltage.
After M05, AOUT1 and AOUT2 should return close to 0 V. If voltage remains, the inverter may continue receiving a speed reference.

Fifth, confirm the I/O mapping.
If the Standard PLC Configuration assigns spindle forward or spindle enable to another output point, but the technician is checking Y1.0, the fault diagnosis will be wrong from the beginning.


11. Correct Principles for Modifying Spindle I/O Points

If an output point is confirmed faulty, such as Y1.0 leakage or no output, the standard PLC configuration can be used to move the function to a spare output point. The following principles should be followed.

First, back up all PLC files.
Before any modification, back up:

PLC210B.COM
PLC210B.CPP
PLCTAB_M.DAT
PLC_MAP.H
PMESSAGE.TXT
related parameter files

Do not use the Load function before a backup is completed, because it may overwrite the active PLC logic.

Second, record the current spindle point assignment.
In the External I/O Point Definition table, record the input/output group, bit, and valid level for:

Spindle forward.
Spindle reverse.
Spindle enable.
Spindle stop.
Spindle orientation.
Spindle zero speed.
Spindle speed reached.

Third, choose a true spare output point.
The spare point must satisfy these conditions:

It is not used by another function in the configuration table.
It has a physical terminal available for wiring.
Its output type matches the original output.
The load current does not exceed the CNC output capacity.
It uses the correct common reference and power supply.

Fourth, modify group, bit, and valid level carefully.
For example, if spindle forward is moved from Y1.0 to Y2.3, change the output group to 2 and the bit to 3. The valid level should usually remain the same as the original logic. If the valid level is changed incorrectly, spindle action may become inverted, which is dangerous.

Fifth, move the external wiring at the same time.
The relay coil or inverter start signal originally connected to Y1.0 must be moved to the new physical output terminal. The software assignment and physical wiring must match.

Sixth, save, reboot, and retest.
After saving the configuration, the CNC or PLC configuration usually needs to be restarted or reloaded. After reboot, test emergency stop, alarms, Y status, XS91 or the new output terminal, relay action, analog voltage, and inverter operation again.


12. Likely Causes of the Customer’s Original Fault

The customer described the fault as follows:

A certain point seems to have a signal going out.
The relay remains energized.
The system screen shows low level.
The main controller display does not match the actual signal.
After the panel stops spindle rotation, the spindle still rotates slowly.
A new inverter has been installed, and the inverter appears normal.

Based on this diagnosis, several conclusions can be made.

First, the inverter should not be the only suspect.
If the spindle continues to rotate after stop, the cause may be that the start relay has not dropped out, or that the analog speed reference has not returned to zero. Replacing the inverter does not eliminate the CNC output circuit or external relay circuit.

Second, if the system shows low level but the relay remains energized, three causes are most likely.
The CNC output transistor or optocoupler may be leaking or shorted.
The external relay circuit may be backfeeding voltage into the output.
The wrong Y point may be checked, while the relay is actually controlled by another point assigned in the standard PLC configuration.

Third, low-speed spindle crawling requires analog voltage measurement.
If AOUT1 or AOUT2 still has voltage after M05, the inverter may interpret it as a speed command. If the run command is also not fully removed, the spindle may rotate slowly.

Fourth, I/O point modification on this HNC-210B should focus on the External I/O Point Definition table.
This system may not provide a direct ladder diagram view. Instead, it may define spindle, axis ready, alarm, tool magazine, cooling, lubrication and other machine functions through the standard PLC configuration table. The customer’s request to “modify the point” is probably a request to modify the group, bit and valid level in this table, not necessarily to rewrite a complete ladder program.


13. Recommended Full Repair Procedure

For an HNC-210B spindle output fault, the repair should follow this sequence.

  1. Record machine model, software version, and PLC file list.
    Photograph the nameplate, software versions, PLC file names, and modification dates.
  2. Clear the bench emergency stop condition.
    Check XT1 emergency stop contacts, XS8 handheld emergency stop, XS10 external run permission input, and panel emergency stop release.
  3. Read the active alarm list.
    Record X/Y/Z drive not ready, CNC alarm, emergency stop, spindle alarm, or other active alarms.
  4. Enter M03 S100 in MDI.
    Confirm whether the screen shows M=3 and S=100.
  5. Check F status.
    Focus on F06.4, F06.7, and F07.4. If F06.7 remains 0, the CNC has not sent the spindle forward request to the PMC.
  6. Check G status.
    Focus on G01.2. If G01.2 is 1, spindle lock is active. If it is 0, spindle lock can be excluded.
  7. Check Y status.
    Focus on Y[001], especially Y1.0, Y1.1, and Y1.2.
  8. Measure XS91 directly.
    Use LED test lamps for Y1.0 and Y1.2. Use a multimeter for AOUT1 and AOUT2.
  9. Enter Setting → PLC → HIG.
    Check the system disk files such as PLC210B.COM, PLC210B.CPP, PLCTAB_M.DAT, PLC_MAP.H, and PMESSAGE.TXT.
  10. Enter the Standard PLC Configuration System.
    Find External I/O Point Definition and photograph the spindle forward, reverse, enable, stop, orientation, zero speed, and speed-reached assignments.
  11. Decide whether point reassignment is necessary.
    If the original output is confirmed faulty, select a spare output point and modify both the software assignment and external wiring.
  12. Save, restart, and retest.
    After modification, retest M03, M05, LED output, analog output, relay action, inverter run command, and zero-speed stop behavior.

14. Maintenance Precautions

Do not randomly load PLC files.
PLC210B.COM may be the active runtime file. Loading the wrong file may overwrite the current machine logic.

Do not casually modify the standard PLC configuration.
Spindle, tool magazine, axis ready, limit, reference return, emergency stop, hydraulic pressure, lubrication, and other functions may be affected.

Do not short analog ground to digital ground randomly.
AOUT AGND should be used only as the analog reference unless the original wiring design proves otherwise.

Do not drive large loads directly with CNC outputs.
CNC Y outputs should normally drive small relays or input circuits. Contactors and high-current loads require intermediate relays or isolation modules.

Do not treat bench short circuits as final repairs.
X1.7, emergency stop, and handheld emergency stop shorts are only for diagnosis. The original safety chain must be restored before the machine is returned to operation.

Update the wiring diagram after I/O point changes.
If the point assignment is changed but the documentation is not updated, future maintenance will become difficult and misdiagnosis will be likely.


15. Conclusion

For HNC-210B spindle faults such as spindle not starting, spindle crawling after stop, or a relay remaining energized while the diagnostic screen shows low level, the problem must not be treated simply as an inverter fault or CNC board fault. The correct method is to divide the control chain into six layers:

CNC command state.
F/G signal status.
PMC/PLC point logic.
Y output status.
XS91 physical interface measurement.
External electrical circuit and inverter terminals.

In this case, M03 S100 was accepted by the CNC screen, but F06.7 did not become active and F07.4 remained active. This means the normal spindle start path was blocked by CNC alarm or machine interlock. G01.2 was 0, so spindle lock was excluded. Therefore, the next diagnostic focus should be:

Direct XS91 measurement of Y1.0, Y1.2, AOUT1, and AOUT2.
Standard PLC Configuration System inspection of external I/O point definitions.
Confirmation of which output group and bit are assigned to spindle forward, spindle reverse, and spindle enable.

If an output point is confirmed faulty, the spindle function may be reassigned to a spare Y point through the standard PLC configuration table, provided the external wiring is changed accordingly. Before any change, the PLC files and configuration must be backed up. After modification, the CNC must be restarted and the spindle start, stop, relay action, analog zero return, and inverter operation must be tested again.

The key to repairing old integrated CNC systems like the HNC-210B is not replacing boards blindly. The essential skill is understanding how CNC commands, PMC/PLC signals, I/O point definitions, and physical interfaces interact. Once F, G, X, Y, XS91, and the standard PLC configuration table are understood, faults such as “spindle does not start,” “spindle cannot stop,” and “displayed output does not match actual relay state” can be reduced to measurable and repairable fault points.

Posted on

GSK DAD08A Dual-Axis Servo Drive Manual Guide: Dual-Axis Wiring, Single-Axis JOG, Electronic Gear, Parameter Backup, Alarm Isolation and Repair

GSK DAD08A dual-axis servo wiring and axis isolation

GSK DAD08A Dual-Axis Servo Drive Manual Guide: Dual-Axis Wiring, Single-Axis JOG, Electronic Gear, Parameter Backup, Alarm Isolation and Repair

DAD08A Service Focus

GSK DAD08A dual-axis servo wiring and axis isolation

GSK DAD08A is a dual-axis servo drive used on CNC machines, feeding units and compact automation systems. Unlike single-axis DA98 drives, one drive manages two servo motors. The axes may share power bus, cooling and protection, but command signals, encoder feedback, electronic gear and tuning must be handled per axis.

Do not judge the whole drive as failed before identifying whether the alarm belongs to axis 1, axis 2 or a shared power/cooling section.

Wiring Check

Check main power, grounding, braking resistor, fan and cooling path first. Label motor cables, brake cables and encoder cables as axis 1 and axis 2. If motor or encoder plugs are swapped, one axis may run while the other alarms, or both directions become confusing.

Return ALM, SRDY and COIN signals separately to CNC or PLC whenever possible. A single common alarm makes troubleshooting slower.

Single-Axis JOG First

Do not run coordinated programs first. Move axes to a safe position, enable axis 1 only, run low-speed JOG or low-frequency pulse, then test axis 2 in the same way. Only after both axes run stably should coordinated motion be restored.

If one axis alarms, swap motor and encoder cables carefully for a short low-speed test. If the fault follows the motor, check motor and encoder. If it stays on the same channel, check the drive channel, terminals or parameters.

Electronic Gear and Parameters

GSK DAD08A dual-axis alarm isolation and parameter backup

Each axis may have different screw pitch, reduction ratio and mechanical direction. Calculate electronic gear separately for each axis. Use a fixed pulse test to measure actual movement. Proportional size error points to electronic gear. Bidirectional error points to backlash or coupling.

Record parameters per axis: control mode, pulse format, direction, electronic gear, speed limit, ramp time, position gain, speed gain, torque limit, in-position window and alarm output logic. Execute parameter write/save before delivery.

Alarm Isolation

Single-axis faults include encoder open circuit, Z phase error, UVW feedback error, following error, overload, direction error and in-position failure. Check the matching encoder cable, motor cable, load and axis parameters.

Shared faults include main power abnormality, DC bus over/undervoltage, cooling fault, braking resistor fault and control power fault. These usually stop both axes.

Coordinated-motion faults occur when single-axis tests are normal but synchronized motion alarms. Check controller pulse frequency, acceleration, mechanical interference and load inertia.

Delivery Flow

Record parameters and wire numbers, confirm motor/encoder matching, check shared power and cooling, JOG axis 1, JOG axis 2, set electronic gear per axis, tune gains per axis, run coordinated low-speed test, verify ALM/SRDY/COIN returns, save parameters and run full load cycle.

Posted on

Daikin Super Unit Manual Guide: Installation, Hydraulic Piping, Electrical Wiring, Filter, Regenerative Resistor, Pressure Sensor, Pump Replacement and Maintenance Diagnosis

Daikin Super Unit installation piping and wiring inspection

Daikin Super Unit Manual Guide: Installation, Hydraulic Piping, Electrical Wiring, Filter, Regenerative Resistor, Pressure Sensor, Pump Replacement and Maintenance Diagnosis

Why Installation and Maintenance Matter

Daikin Super Unit installation piping and wiring inspection

Many Daikin Super Unit faults do not come from the controller board itself. Common causes include poor cooling space, bad suction piping, weak grounding, incorrect filter wiring, pressure sensor cable vibration, hot regenerative resistor and missing calibration after pump replacement. This guide focuses on installation, wiring, maintenance and recalibration.

Mounting, Cooling and Fixing

Leave sufficient cooling space around the controller and motor pump. The manual recommends at least 100 mm clearance. Do not block the motor cooling fan. The motor pump must be fixed firmly because hydraulic reaction force and startup shock can move it.

If the suction port direction is changed, follow the manual procedure. Do not hammer the pump shaft or force the port; damaged spline, shaft or seal surface will later cause noise, leakage, heat and reduced flow.

Hydraulic Piping

The Super Unit does not include a safety valve. Install one on the machine side and set it around maximum operating pressure plus 2 MPa. Avoid inappropriate straight check valves on the discharge side because they may prevent controlled pressure release. Accumulator circuits need backflow protection to prevent reverse pump rotation after power off.

Install a suction filter of about 150 mesh and keep suction pressure within -0.02 MPa. Low suction pressure causes air intake, cavitation noise, reduced flow and pressure pulsation.

Electrical Wiring

Main terminals include L1/L2/L3 power input, U/V/W motor output, P1/P2 DC reactor and B1/B2 regenerative resistor. Use correct wire size, crimp terminals and torque. Power cable and motor cable must never be connected to I/O terminals.

I/O terminals handle analog commands, monitor outputs, alarm contacts, digital I/O, pressure sensor and encoder signals. Keep analog and sensor cables away from U/V/W and contactor coils. Ground shields correctly.

Filter, Grounding and Noise Control

Use an input filter to suppress noise. The filter body must have good metal contact with the cabinet. Keep input and output cables separated; do not bundle them or route them in the same duct. Poor grounding or parallel filter wiring can cause random alarms, monitor noise and PLC input errors.

Regenerative Resistor and DC Reactor

Daikin Super Unit pressure sensor pump replacement and regeneration service

P1/P2 are used for the DC reactor. B1/B2 are used for the regenerative resistor. Regenerative energy appears during deceleration, load-driven rotation or hydraulic backflow. Install the resistor on a ventilated metal surface and prevent accidental touch. Wrong resistor value, capacity or wiring can lead to overvoltage or controller damage.

Pressure Sensor and H30

The pressure sensor is the core of pressure feedback. Install the ferrite core near the controller side of the sensor cable, around 100 mm from the terminal, with two turns. Fix the cable near the sensor so pump vibration does not pull the connector.

After replacing the pressure sensor, set H30 PS_G. Power up after pressure is released so zero point can be corrected, then calibrate with an accurate pressure gauge.

Pump Replacement, H15 and P07

After pump replacement, H15 Q_EV or P07 QMAX may need adjustment. Use monitor mode to observe Qo and compare with metering action, cylinder speed or measured flow. Do not blame parameters before checking suction pipe, filter, oil condition, leakage and valve resistance.

Preventive Service

Record oil temperature, noise, suction filter condition, fan status, Po/Qo monitor values, alarm history and high-pressure duty time. High oil temperature, dirty oil, suction leak and blocked filter often look like parameter problems but are hydraulic-condition problems.

Recommended delivery order: fixing and cooling, safety valve, suction pressure, filter grounding, main terminals, I/O analog signals, pressure sensor fixing, regenerative resistor cooling, low-pressure bleeding, PMAX/QMAX calibration, H30/H15 check and full-cycle machine test.

Posted on

GSK DA98A/DA98D AC Servo Drive Manual Guide: Replacement Matching, Panel JOG, Parameter Save, Electronic Gear, Encoder and Alarm Repair

GSK DA98A DA98D servo replacement and startup JOG flow

GSK DA98A/DA98D AC Servo Drive Manual Guide: Replacement Matching, Panel JOG, Parameter Save, Electronic Gear, Encoder and Alarm Repair

DA98A/DA98D Maintenance Focus

GSK DA98A DA98D servo replacement and startup JOG flow

GSK DA98A and DA98D are later members of the DA98 AC servo drive family. They are used on CNC lathes, milling machines, grinders, feed axes and automatic positioning mechanisms. They share the same service logic as DA98, but replacement requires more attention to motor model, encoder type, firmware/version differences and EEPROM data.

Do not replace only by power rating. Record original parameters and motor nameplate first. The common mistakes are wrong motor-encoder matching, forgetting EEPROM write after parameter changes and misjudging mechanical jam or encoder contact failure as drive failure.

Panel Startup and JOG

Before power on, check main power, grounding, U/V/W motor cable, braking resistor, CN1 control signals, CN2 encoder cable and shield. On first power up, do not enable SON immediately. Check display and alarm code first.

If there is no alarm, use JOG or low-speed trial run with the axis in a safe position. Verify direction, noise, current and encoder feedback. If the motor vibrates or alarms immediately, check phase order, encoder cable, coupling and load inertia before tuning gains.

CN1 Control Logic

CN1 usually exchanges SON enable, ALRS reset, drive inhibit, deviation clear, PULS/SIGN or dual pulse input, SRDY ready, ALM alarm and COIN in-position signals. At minimum, ALM and SRDY should return to CNC or PLC so the controller knows whether the servo is really ready.

If the axis does not move, check CNC pulse output, SON state, ALM, drive inhibit, pulse inhibit, electronic gear ratio and COIN conditions. If direction is reversed, change one location only; do not change CNC parameter, motor phase and drive direction at the same time.

CN2 Encoder and Motor Matching

GSK DA98A DA98D parameter save electronic gear and alarm diagnosis

CN2 carries encoder feedback. Encoder alarms, Z phase loss, UVW error and count error often come from loose connector, backed-out pin, broken shield, 5 V drop or encoder failure. Check cable and feedback first before replacing the drive.

After replacing drive or motor, confirm motor model parameter, encoder line count and feedback direction. Wrong matching may still allow JOG but cause positioning error, crawling, overcurrent or overheating.

Electronic Gear and Position Control

In position mode, electronic gear ratio defines mechanical movement per CNC pulse. No.12 and No.13 are commonly used as numerator and denominator, while No.14 and No.15 set pulse format and direction. Proportional size error points to electronic gear error. Bidirectional inconsistency points to backlash or mechanical issues.

Use a fixed low-frequency pulse test, measure actual movement, calculate No.12/No.13 and then verify with low-speed G01/G00 movement. Do not hide electronic gear error by changing tool offset or error window.

EEPROM Parameter Save

After changing control mode, electronic gear, pulse direction, JOG speed, speed/position gains, torque limit and related parameters, execute EE-SEt to write EEPROM. Otherwise the drive may revert after power-off.

Before replacement delivery, record control mode, pulse input type, electronic gear, maximum speed, acceleration/deceleration time, speed PI, position gain, feedforward, torque limit, alarm output logic and in-position window.

Speed Mode and Torque Limit

Some DA98A/DA98D applications use speed mode rather than pulse position mode. Check mode, speed command, forward/reverse start, maximum speed, ramp time and speed-arrival output. For torque-limit issues, check external limit input, parameter limit, load and motor current.

If no-load test is normal but load operation alarms, check guide lubrication, ball screw bearing, coupling alignment and load inertia. Parameters cannot repair a jammed machine.

Alarm Layers

Power and main-circuit alarms: check input voltage, braking resistor, DC bus capacitor, power module, contactor and grounding. Overvoltage during deceleration often relates to braking energy, inertia or too short deceleration time.

Encoder alarms: check CN2 cable, connector, shield, 5 V supply and encoder before replacing the drive.

Following error: check electronic gear ratio, command frequency, position gain, mechanical jam, backlash and inertia.

Overcurrent, overload and overheating: check motor phase cable, load jam, ramp time, speed gain, fan cooling and motor insulation. Hard alarms that cannot be cleared by ALRS require power-off inspection.

Recommended Commissioning Flow

Record original parameters and nameplates, confirm motor and encoder matching, power up without alarm, use low-speed JOG, wire CN1 basic signals, calculate electronic gear, test low-frequency pulse movement, tune gains and torque limits, execute EE-SEt and finally verify positioning accuracy under load.

Posted on

Daikin SUT Injection Molding Machine Super Unit Manual Guide: Pi/Qi Pressure and Flow Commands, I/O Wiring, PMAX/QMAX, Response Parameters, Bleeding and Alarm Repair

Daikin SUT Super Unit Pi Qi pressure and flow control wiring

Daikin SUT Injection Molding Machine Super Unit Manual Guide: Pi/Qi Pressure and Flow Commands, I/O Wiring, PMAX/QMAX, Response Parameters, Bleeding and Alarm Repair

SUT Is a Servo Hydraulic Power Unit, Not a Simple Inverter Pump

Daikin SUT Super Unit Pi Qi pressure and flow control wiring

The Daikin SUT Super Unit for injection molding machines is a variable-speed IPM motor hydraulic system. The controller, motor pump, encoder, pressure sensor and machine hydraulic circuit form pressure and flow feedback loops. The injection molding controller sends analog command voltages; the SUT controls pump displacement Q by motor speed and controls pressure P by pressure sensor feedback.

For repair, do not treat it like a normal inverter pump. Check Pi pressure command, Qi flow command, Po pressure monitor, Qo flow monitor, pressure sensor, encoder, piping, oil condition, overload duty and valve timing together.

Matching Controller and Motor Pump

The manual states that the motor pump and controller are factory tested as a matched pair. Confirm serial numbers and model compatibility before replacement. When replacing the controller, pump or pressure sensor, check H15 flow correction, H30 pressure sensor gain and P07 QMAX. These values strongly affect actual hydraulic output.

Hydraulic Piping and Oil Conditions

The SUT does not include a safety valve. Install a safety valve on the injection molding machine side, normally set to maximum operating pressure plus 2 MPa. Use a suction filter of about 150 mesh and keep suction pressure within -0.02 MPa. Low suction pressure causes air intake and cavitation noise, with higher noise, reduced flow and possible pump damage.

After installation or repair, bleed the hydraulic circuit by moving each cylinder. Long high-pressure manual operation may cause E27 or E17 overload alarms; power off the controller to reset and continue under non-overload conditions.

Main Circuit and I/O Wiring

Main terminals include L1/L2/L3 power input, P1/P2 DC reactor, B1/B2 regenerative resistor and U/V/W motor output. Power, motor and signal wiring must not be mixed. Grounding must be reliable.

The I/O terminal block includes AI1 Pi pressure command, AI2 Qi flow command, AO1 Po pressure monitor, AO2 Qo flow monitor, AGND, pressure sensor signal and supply, ALM_A/ALM_B/ALM_C alarm output, DICOM/DOCOM and digital I/O. Separate analog ground, digital ground, shield and frame ground carefully.

The pressure sensor cable should use the specified connector. The ferrite core is fitted near the controller side, around 100 mm from the terminal, with two turns. Fix the cable near the sensor to prevent pump vibration from damaging the connector.

Pi/Qi and Po/Qo Linearity

Pi is pressure command voltage and Qi is flow command voltage, usually DC 0-10 V corresponding to 0-PMAX and 0-QMAX. Po and Qo are monitor outputs. Pi/Qi must be linear with the injection molding controller’s 0-99.9% pressure and flow settings. Otherwise the machine panel setting cannot accurately control hydraulic output.

If maximum analog voltage is below 10 V, set P05 VMAX and then calibrate PMAX/QMAX accordingly. During service, measure AI1/AI2 and AO1/AO2 at 20%, 50% and 80% settings.

Initial Parameters

Daikin SUT response parameters alarms and maintenance diagnosis

P06 is PMAX maximum control pressure. P07 is QMAX maximum discharge flow. The manual emphasizes that P06 and P07 must be set before operation. P07 is not theoretical pump displacement; it is a calibrated value based on actual pump speed and flow. P05 is the common command voltage scale for pressure and flow. WN_L is the overload warning output threshold.

If the same machine setting produces different speed or pressure after replacement, check P06, P07, P05, H15 and H30 before replacing hardware again.

Response Parameters

P08 P_UG adjusts pressure rise response. Higher value gives faster response but more overshoot. P09 P_DG adjusts pressure fall response. P10 and P11 adjust flow rise and fall response. P13 SC_G reduces pressure overshoot. P14 D_TM delays response from standby, allowing hydraulic valves to switch before the pump builds pressure.

P17-P21 are pressure proportional gain and integral time parameters. P22 and P23 set pressure rise/fall ramp time. Tune in this order: oil and piping first, Pi/Qi linearity second, P08-P11 basic response third, then P13/P14/P22/P23 for shock and overshoot control.

Startup and Bleeding

Do not load the pump within five seconds after power on. During bleeding, use manual low-risk operation and observe pressure, flow and noise. Crawling during injection unit movement may be caused by air in the circuit or pressure overshoot. Use moderate pressure and flow-control operation to verify stability.

Alarm Repair

E17/E27 overload alarms are often related to high pressure and high flow duty, long high-pressure manual bleeding, abnormal oil temperature, excessive load, fan trouble or poor cooling. Reduce duty, check oil temperature and cooling, then review aggressive response parameters.

Pressure sensor faults may come from open circuit, short circuit, incorrect wiring, abnormal pressure, poor connector contact or sensor failure. Check supply, signal, AGND, ferrite core and connector before replacing the sensor. After replacement, set H30 PS_G and calibrate with an accurate pressure gauge.

Low flow and loud noise usually point to suction pressure, suction filter, oil viscosity, oil temperature, suction leak or cavitation. Do not hide suction problems by increasing Q_UG or QMAX.

After pump replacement, check P07 QMAX and H15 Q_EV. In monitor mode, verify Qo and compare actual cylinder speed or metering flow.

Service Delivery Checklist

Before delivery, record P05, P06, P07, P08-P11, P13, P14, P16, P17-P23, WN_L, H15 and H30. Also record Pi/Qi, Po/Qo, current, oil temperature and alarm history during mold open/close, injection, holding, metering and ejection.

The reliable SUT repair sequence is: confirm pump-controller matching, check hydraulic piping and safety valve, verify I/O and analog linearity, set PMAX/QMAX, bleed the circuit, tune response parameters carefully and then diagnose overload, sensor or oil-circuit faults from alarm history.