About the GALAD Score for Hepatocellular Carcinoma Detection
Five inputs into one logistic model: Z = −10.08 + 0.09 × age + 1.67 × sex + 2.34 × log₁₀(AFP) + 0.04 × AFP-L3% + 1.33 × log₁₀(DCP), with sex coded 1 for male and 0 for female. The commonly applied threshold is −0.63, at which the development and validation work reported roughly 86% sensitivity and 90% specificity for hepatocellular carcinoma including early-stage disease. Note which terms are transformed: AFP is log₁₀ with a coefficient of 2.34, while AFP-L3 enters as a raw percentage with a coefficient of 0.04. Those two are frequently quoted the wrong way round, and swapping them changes the score substantially.
Formula
Z = -10.08
+ 0.09 x age (years)
+ 1.67 x sex (1 = male, 0 = female)
+ 2.34 x log10(AFP) (ng/mL)
+ 0.04 x AFP-L3 (raw %, NOT logged)
+ 1.33 x log10(DCP) (ng/mL)
Threshold commonly applied: Z >= -0.63 is suspicious for HCC
(approximately 86% sensitivity, 90% specificity)- 2.34 x log₁₀(AFP)
- The dominant biomarker term. Each tenfold rise in AFP adds 2.34 to the score, so AFP moving from 10 to 1000 ng/mL contributes 4.68 on its own.
- 0.04 x AFP-L3
- Raw percentage, small coefficient. Even a markedly abnormal AFP-L3 of 50% adds only 2.0 — the opposite of what the acronym's prominence suggests.
- 1.33 x log₁₀(DCP)
- The second logged biomarker. Sensitive to units: an assay reporting mAU/mL instead of ng/mL will not give the score the model was fitted on.
- 1.67 x sex
- A flat male increment, and a large one — bigger than the contribution of a tenfold change in DCP. It reflects the substantially higher baseline incidence in men.
- AFP and DCP are log₁₀-transformed; age and AFP-L3 enter untransformed.
- AFP and DCP must both exceed zero, since log₁₀(0) is undefined.
- The score is unbounded and can be strongly negative; it is not a percentage or a probability.
- It was developed in patients with chronic liver disease under surveillance, and its performance outside that population is not established.
Interpreting the result
A score at or above −0.63 should trigger diagnostic imaging, not a diagnosis: arrange multiphase contrast-enhanced CT or MRI, and let LI-RADS characterise whatever is found. Before acting on a high score, confirm the DCP units, since a PIVKA-II result in mAU/mL entered as ng/mL is the commonest way to inflate it. Check also whether the patient is on warfarin or vitamin K deficient — both raise DCP through the same mechanism as tumour does and will produce a false positive. A score below the threshold does not exclude cancer and should not alter the surveillance interval; six-monthly ultrasound continues regardless. Trend is more informative than any single value, and a score rising steadily within the negative range deserves attention that a static one does not. The most important boundary is the population: GALAD was developed and validated in patients with chronic liver disease under surveillance, and applying it to someone without that background risk inflates the meaning of every value it returns.
| Score | Band | What it means | Action |
|---|---|---|---|
| Z ≥ −0.63 | Suspicious for hepatocellular carcinoma | At the commonly applied operating point — roughly 86% sensitivity, 90% specificity | Arrange multiphase contrast-enhanced CT or MRI; confirm DCP units and exclude warfarin or vitamin K deficiency first |
| Z < −0.63 | Below threshold | Does not exclude hepatocellular carcinoma | Continue six-monthly surveillance ultrasound; watch the trend rather than the single value |
What the GALAD Score needs (5 inputs)
- Age
- In years, contributing 0.09 per year. Over a plausible surveillance age range this term spans several points and is the largest non-biomarker contributor.
- Sex
- Coded 1 for male and 0 for female, contributing a flat 1.67 when male. That single term is larger than a tenfold difference in DCP.
- AFP
- In ng/mL, log₁₀-transformed, coefficient 2.34. This is the term carrying the large coefficient — not AFP-L3. Because it is logged it must be above zero.
- AFP-L3
- As a raw percentage, NOT log-transformed, coefficient 0.04. A 50% AFP-L3 therefore adds only 2.0 to the score. Treating this as the logged term is the commonest transcription error.
- DCP (des-gamma-carboxy prothrombin)
- In ng/mL, log₁₀-transformed, coefficient 1.33. Also called PIVKA-II. Many laboratories report the same analyte in mAU/mL, and substituting that figure without conversion shifts the result.
What it returns
- GALAD score (Z)
- A continuous logistic score. Negative values are common in patients without cancer; the scale is not a probability.
- Position relative to the −0.63 threshold
- The most widely applied operating point. Other cut-offs appear in the literature depending on whether sensitivity or specificity is being favoured.
- The individual term contributions
- Broken out so it is visible which variable is driving the score — usually AFP, occasionally the sex term alone.
How it is calculated
The three biomarkers in GALAD fail in different patients, which is the whole point of combining them. AFP is raised in a majority of large tumours but in a minority of small ones, and it also rises with hepatic inflammation, so it is neither sensitive early nor specific in active hepatitis. AFP-L3 is the fucosylated fraction of AFP produced preferentially by malignant hepatocytes, so it improves specificity — but it is only interpretable when there is enough AFP to fractionate, which limits it in exactly the low-AFP tumours where help is most needed. DCP is an abnormal prothrombin produced when tumour cells lose vitamin K-dependent carboxylation, and it is largely independent of the other two, rising in a partly different set of tumours. Because their failures are uncorrelated, a logistic model over all three detects tumours that no single marker would. Adding age and sex then adjusts for the fact that the same biomarker profile means something different in a 45-year-old woman and a 70-year-old man — which is why the sex term is as large as it is.
Facts & figures
| Term | Transformation | Coefficient | Worked example |
|---|---|---|---|
| Age | None (years) | 0.09 | Age 60 adds 5.40 |
| Sex | 1 male / 0 female | 1.67 | Male adds 1.67; female adds 0 |
| AFP | log₁₀ (ng/mL) | 2.34 | AFP 100 adds 4.68; AFP 1000 adds 7.02 |
| AFP-L3 | None (raw %) | 0.04 | AFP-L3 50% adds only 2.00 |
| DCP | log₁₀ (ng/mL) | 1.33 | DCP 10 adds 1.33 |
| Constant | — | −10.08 | Applied to every patient |
The AFP and AFP-L3 rows are the pair most often transposed. AFP is the logged term with the large coefficient; AFP-L3 is the raw term with the small one.
| Marker | What it reflects | Where it fails |
|---|---|---|
| AFP | Bulk tumour production of a foetal protein | Insensitive in small tumours; raised by hepatic inflammation |
| AFP-L3 | The fucosylated fraction made preferentially by malignant hepatocytes | Needs enough total AFP to fractionate, so weak in low-AFP disease |
| DCP | Abnormal prothrombin from loss of vitamin K-dependent carboxylation | Raised by warfarin and vitamin K deficiency regardless of tumour |
Their failure modes are largely independent, which is why combining them detects tumours that none of them finds alone.
Evidence
Derivation and validation — Johnson and colleagues, 2014
2014A prospectively developed statistical model built on serological biomarkers in patients with chronic liver disease, published in Cancer Epidemiology, Biomarkers and Prevention in 2014, with validation in independent cohorts.
At the −0.63 operating point the model reported approximately 86% sensitivity and 90% specificity for hepatocellular carcinoma, with performance retained in early-stage disease — the property that distinguishes it from AFP alone.
International validation — GALAD and BALAD-2
2016Berhane and colleagues assessed the GALAD model for diagnosis and BALAD-2 for survival prediction across international cohorts, reported in Clinical Gastroenterology and Hepatology in 2016.
GALAD retained discrimination for hepatocellular carcinoma across geographically distinct populations, and outperformed any of its individual biomarker components.
Guideline context — AASLD 2018
2018AASLD practice guidance on the diagnosis, staging and management of hepatocellular carcinoma.
Frames surveillance around six-monthly ultrasound with or without AFP; GALAD is an adjunct to that pathway rather than a replacement for it, and its biomarker components are not universally available.
How it compares
GALAD Score vs CT/MRI LI-RADS v2018
Sequential, not competing — GALAD decides who to image, LI-RADS decides what the image shows.
GALAD operates on blood in a patient with no known lesion and answers whether the suspicion is high enough to warrant cross-sectional imaging. LI-RADS operates on that imaging once an observation exists, and assigns a category from enhancement pattern and size. They share a population — both are defined only for patients at risk of hepatocellular carcinoma — and neither is meaningful outside it. A raised GALAD with an LR-5 observation is a diagnosis; a raised GALAD with no imaging correlate is a reason to shorten the interval and repeat, not to keep escalating the biomarkers.
GALAD Score vs Alpha-fetoprotein alone
The comparison GALAD exists to win, and the margin is largest exactly where it matters — early-stage disease.
AFP alone performs poorly enough as a surveillance test that several guidelines stopped recommending it in isolation: it is insensitive in small tumours and non-specific in active hepatitis, so it both misses cancers and generates work. GALAD keeps AFP but logs it, adds two markers whose failure modes are different, and adjusts for the demographic variables that shift baseline risk. The reported gain is a sensitivity around 86% at 90% specificity with performance retained in early-stage disease. The practical cost is that AFP-L3 and DCP are not routinely available in many laboratories, which is the main reason the score is not more widely used.
GALAD Score vs Milan criteria
Opposite ends of the same patient journey — one asks whether there is a cancer, the other what can be done about it.
GALAD is a detection model applied when nothing has been found, using blood alone. Milan is an allocation rule applied once a tumour is confirmed and staged, using the number and size of lesions plus the absence of vascular invasion and extrahepatic spread. Nothing in GALAD informs transplant eligibility and nothing in Milan informs detection. They belong together only in the sense that effective surveillance is what puts patients inside Milan in the first place — the argument for a sensitive early-detection test is that it finds tumours while they are still transplantable.
Pearls & pitfalls
- AFP is the log₁₀ term with coefficient 2.34; AFP-L3 is the raw percentage with coefficient 0.04. Transposing them is the classic error and produces plausible but wrong scores.
- Confirm DCP is reported in ng/mL. The same analyte is widely reported as PIVKA-II in mAU/mL, and substituting it without conversion invalidates the score.
- Warfarin and vitamin K deficiency raise DCP by the same mechanism as tumour, and are a genuine cause of false positives.
- AFP and DCP are logged, so neither can be zero — a reported result of zero needs the assay's lower limit substituted, not a literal 0.
- The sex term is a flat 1.67 for males, larger than a tenfold change in DCP.
- A negative score does not exclude cancer and must not lengthen the surveillance interval.
- The trend across serial measurements is more useful than any single value.
- It is a detection model for patients already under surveillance — not a screening test for the general population, and not prognostic.
- AFP-L3 is unreliable when total AFP is very low, which is exactly the situation where extra sensitivity would help most.
- A raised score is an indication to image, and LI-RADS then does the characterising — the two tools sit in sequence.
Critical actions
- Confirm the patient is in a surveillance population — cirrhosis or chronic hepatitis B — before applying the score at all.
- Check the DCP assay units and convert if the laboratory reports PIVKA-II in mAU/mL.
- Review anticoagulation and nutritional status before accepting a raised DCP.
- Substitute the assay's lower limit of detection where AFP or DCP is reported as zero.
- Arrange multiphase contrast-enhanced CT or MRI for a score at or above the threshold.
- Continue six-monthly ultrasound irrespective of a below-threshold result.
- Record the score alongside its date so the trend can be read at the next visit.
- Do not use the score to stage or prognosticate a tumour that has already been found.
Why this score exists
The most instructive thing about GALAD is how small the AFP-L3 coefficient is, because it contradicts the intuition the acronym creates. AFP-L3 is the most specialised assay of the three, the hardest to obtain, and the one that sounds most sophisticated — and it carries a coefficient of 0.04 against a raw percentage, so a wildly abnormal 50% moves the score by 2.0. AFP, the ordinary test everyone already has, is logged and carries 2.34, so a single order of magnitude in AFP outweighs the entire plausible range of AFP-L3. That ordering is not an oversight; it is what the regression found. It also explains why the two are so often transposed in secondary descriptions: people reproduce the formula from memory and assume the exotic marker must be doing the heavy lifting. It is worth checking any implementation against the original on exactly this point, because a swapped pair still produces plausible-looking numbers and will never announce itself.
About the creator
First author; hepatocellular carcinoma biomarkers and prognostic modelling
Led the development and validation of the GALAD and BALAD family of serological models.
Biostatistician; GALAD and BALAD-2 validation
Led the international validation of the model across geographically distinct cohorts.
Co-author; hepatology and HCC biomarker research
Contributed the Japanese cohorts in which AFP-L3 and DCP were already in routine use.
Limitations
- AFP-L3 and DCP are not routinely available in many laboratories, which is the principal barrier to using the score at all.
- DCP is raised by warfarin and vitamin K deficiency independently of tumour, producing false positives in a common patient group.
- Assay units differ between centres, and a PIVKA-II value in mAU/mL substituted for DCP in ng/mL silently changes the result.
- Developed and validated in patients under surveillance for chronic liver disease; performance outside that population is not established.
- AFP-L3 is unreliable at low total AFP, weakening the model in precisely the tumours that are hardest to detect.
- Reported cut-offs vary between studies depending on whether sensitivity or specificity is prioritised, so −0.63 is a convention rather than a fixed property.
- The score is a detection aid only — it carries no staging or prognostic information once a tumour is found.
- Aetiology is not in the model, though the biomarker distribution differs between viral and metabolic liver disease.
If you are the patient
The GALAD score is a blood test result used in people who are already having regular checks for liver cancer — usually because they have cirrhosis or long-standing hepatitis B. It combines three blood markers made by liver tumours (AFP, AFP-L3 and DCP) with your age and sex, and turns them into a single number. The point of combining them is that each marker misses a different group of tumours, so together they find cancers that any one of them alone would not. A score above the cut-off does not mean you have cancer. It means the blood picture is enough of a signal to justify a proper scan — usually a CT or MRI with contrast dye — which is what actually shows whether there is a tumour. Quite often the scan is clear. A score below the cut-off is reassuring but does not rule cancer out, and it does not mean your routine scans can be spaced out; the six-monthly ultrasound continues either way. Two practical things are worth mentioning to your doctor. If you take warfarin, or if you have had poor nutrition or problems absorbing vitamins, one of the three markers can be raised for reasons that have nothing to do with cancer — so tell them, because it changes how the result is read. And these results are most useful compared with your previous ones: a number that is drifting upwards over time says more than a single reading.
Frequently asked questions
What is the GALAD score?#
A logistic model combining Gender, Age, AFP-L3, AFP and DCP to estimate the likelihood that a patient under surveillance for chronic liver disease has hepatocellular carcinoma. The equation is Z = −10.08 + 0.09 × age + 1.67 × sex + 2.34 × log₁₀(AFP) + 0.04 × AFP-L3% + 1.33 × log₁₀(DCP), with sex coded 1 for male and 0 for female.
What is the GALAD cut-off?#
−0.63 is the most widely applied threshold, at which the development and validation work reported roughly 86% sensitivity and 90% specificity, including for early-stage disease. Other cut-offs appear in the literature depending on whether a study is favouring sensitivity or specificity, so the figure is a convention rather than a fixed property of the model.
Which variables are log-transformed?#
AFP and DCP, both base-10. Age and AFP-L3 enter untransformed. This is the detail most often reproduced incorrectly: AFP is the logged term with the large coefficient of 2.34, while AFP-L3 enters as a raw percentage with a coefficient of only 0.04. Swapping the two produces plausible-looking but substantially wrong scores.
Why is the AFP-L3 coefficient so small?#
Because that is what the regression found, and it is counter-intuitive given how specialised the assay is. AFP-L3 enters as a raw percentage, so even a markedly abnormal 50% contributes only 2.0 to the score — whereas a single tenfold rise in AFP contributes 2.34. The ordinary test does more work than the exotic one.
What causes a falsely raised GALAD score?#
Most often DCP. It is an abnormal prothrombin produced when vitamin K-dependent carboxylation fails, so warfarin therapy and vitamin K deficiency both raise it by the same mechanism a tumour does. Unit confusion is the other common cause — the same analyte is widely reported as PIVKA-II in mAU/mL rather than DCP in ng/mL, and substituting one for the other shifts the score.
Does a low GALAD score exclude liver cancer?#
No. It is a detection aid, not a rule-out test, and a below-threshold score should not lengthen the surveillance interval. Six-monthly ultrasound continues regardless. A score that is rising over serial measurements, even within the negative range, is more informative than any single value.
Can GALAD be used for population screening?#
No. It was developed and validated in patients with chronic liver disease already under surveillance, and its performance depends on that elevated baseline risk. Applying it to a general population, where the prevalence of hepatocellular carcinoma is far lower, would inflate the meaning of every positive result it produced.
How does GALAD relate to LI-RADS?#
They run in sequence rather than in competition. GALAD works on blood in a patient with no known lesion and decides whether to obtain cross-sectional imaging; LI-RADS works on that imaging and categorises whatever observation is found. Both are defined only for the at-risk population, and a raised GALAD with no imaging correlate is a reason to repeat sooner, not to escalate biomarker testing further.
References
Original / primary reference
Clinical practice guidelines
- Marrero JA, Kulik LM, Sirlin CB, Zhu AX, Finn RS, Abecassis MM, et al. Diagnosis, Staging, and Management of Hepatocellular Carcinoma: 2018 Practice Guidance by the American Association for the Study of Liver Diseases. Hepatology. 2018;68(2):723-750.
- European Association for the Study of the Liver. EASL Clinical Practice Guidelines: Management of hepatocellular carcinoma. J Hepatol. 2018;69(1):182-236.