About the VOCAL-Penn Cirrhosis Surgical Risk Score
The VOCAL-Penn score estimates the risk of death 30, 90 and 180 days after surgery — and of hepatic decompensation within 90 days — in a patient with cirrhosis. It uses nine pre-operative variables: age, serum albumin, total bilirubin, platelet count, ASA physical status class, obesity, MASLD/NAFLD aetiology, whether the operation is an emergency, and which type of operation is planned. Unlike MELD or Child-Turcotte-Pugh, it accounts for the specific operation being considered, and in the original study it discriminated better than MELD, MELD-Na, Child-Turcotte-Pugh and the Mayo Risk Score at every time point.
Formula
risk = logistic( β₀ + f(albumin) + f(bilirubin) + f(platelets) + β·age + β·ASA class + β·surgery category + β·emergency + β·MASLD + β·obesity )- f( )
- A restricted cubic spline, not a straight line. Albumin, bilirubin and platelet count each enter the model as a spline so their effect on risk bends at clinically meaningful values instead of rising in fixed increments.
- logistic( )
- The logistic function, which turns the linear predictor into a probability between 0 and 1.
- β
- Published regression coefficients from the derivation model, one per variable and one per surgery category relative to laparoscopic abdominal surgery as the reference.
- There is deliberately no arithmetic to reproduce by hand here. Separate models exist for 30-day, 90-day and 180-day mortality and for 90-day decompensation, and the 90- and 180-day models are conditional on surviving the preceding window.
- The calculator combines the conditional models with the 30-day model using the law of total probability, so the 90- and 180-day figures it reports are cumulative rather than conditional.
- This is the reason the score is under-used relative to MELD: MELD can be done on paper, and this cannot.
Interpreting the result
The output is a continuous probability, not a risk band — the authors deliberately publish no low/medium/high cut-offs. Read the 30-day figure alongside the 90- and 180-day figures: a patient whose risk climbs steeply between 30 and 180 days has a liver-disease trajectory that surgery alone will not fix. Compare predictions across operative approaches where a choice exists, and use the 90-day decompensation estimate to plan post-operative monitoring even when mortality risk looks acceptable.
What the VOCAL-Penn Score needs (9 inputs)
- Age
- Age in years at the time of surgery. The model was derived in patients aged roughly 40–85.
- Serum albumin (g/dL)
- Pre-operative albumin. Modelled non-linearly with a restricted cubic spline; lower albumin increases predicted risk.
- Total bilirubin (mg/dL)
- Pre-operative total bilirubin. Higher bilirubin increases predicted 30-day mortality and 90-day decompensation.
- Platelet count (×1,000/µL)
- Serves as the model's marker of portal hypertension. The authors chose it over ascites because ascites assessment is subjective, and the two performed near-interchangeably.
- ASA physical status class
- American Society of Anesthesiologists class, entered as an ordinal 2, 3 or 4. Classes 1, 5 and 6 were excluded when the model was derived.
- Surgery type
- One of six categories: laparoscopic abdominal (the reference), open abdominal, abdominal wall, vascular, major orthopedic, or chest/cardiac.
- Emergency indication
- Whether the operation is being performed as an emergency. Emergency surgery roughly 2.5× the odds of 30-day death in the derivation cohort.
- MASLD / NAFLD cirrhosis
- Whether the aetiology of cirrhosis is metabolic dysfunction-associated steatotic liver disease (formerly NAFLD), which carried a higher post-operative mortality than other aetiologies.
- BMI ≥ 30 kg/m²
- Obesity was independently associated with lower predicted post-operative mortality in this cohort (30-day odds ratio 0.47).
Units. Enter albumin in g/dL, bilirubin in mg/dL and platelets in ×1,000/µL. If your laboratory reports in SI units, convert first: bilirubin µmol/L ÷ 17.1 gives mg/dL, and albumin g/L ÷ 10 gives g/dL. Platelet count is numerically the same whether reported as ×1,000/µL or ×10⁹/L.
What it returns
- 30-day post-operative mortality
- Probability of death within 30 days of the operation.
- 90-day post-operative mortality
- Cumulative probability of death within 90 days, combining the 30-day risk with the conditional risk among 30-day survivors.
- 180-day post-operative mortality
- Cumulative probability of death within 180 days, extending the same conditional logic.
- 90-day hepatic decompensation
- Probability of new ascites, hepatic encephalopathy or variceal bleeding within 90 days of surgery, from the companion VOCAL-Penn decompensation model.
How it is calculated
VOCAL-Penn is not an additive point score. It is a set of multivariable logistic regression models in which albumin, platelet count and bilirubin are entered as restricted cubic splines rather than straight lines, so their effect on risk bends at clinically meaningful thresholds. The 90- and 180-day models are conditional on surviving the previous window, and the calculator combines them with the 30-day model using the law of total probability to produce cumulative risks. Because the arithmetic cannot be done by hand, a calculator is the only practical way to apply the score at the bedside.
Evidence
Derivation — US Veterans Affairs VOCAL cohort
2021 · n = 3,7854,712 surgical procedures in 3,785 patients with cirrhosis, split 80/20 into derivation and internal validation sets. Abdominal, abdominal wall, vascular, major orthopedic and chest/cardiac procedures were all included.
30-day C-statistic 0.859 (95% CI 0.809–0.909) in internal validation, versus 0.766 for the Mayo Risk Score, 0.752 for MELD-Na, 0.724 for MELD and 0.682 for Child-Turcotte-Pugh. Calibration was good at each time point, whereas the Mayo Risk Score overestimated risk across the whole spectrum.
External validation — Beth Israel Deaconess Medical Center and University of Pennsylvania Health System
2021 · n = 855855 surgical procedures between January 2008 and October 2015, in two health systems independent of the derivation cohort and of each other.
Numerically the highest C-statistic for 90-day post-operative mortality at 0.82, against 0.79 for the Mayo Risk Score, 0.79 for MELD and 0.78 for MELD-Na — though these differences did not reach statistical significance in a cohort this size. VOCAL-Penn did have the lowest Brier score and the highest index of prediction accuracy at both time points, which points to better overall model performance rather than discrimination alone.
How it compares
VOCAL-Penn Score vs MELD-Na
Use VOCAL-Penn for the surgical question and MELD-Na for liver disease severity — they answer different questions, and MELD-Na was never built to account for the operation.
MELD-Na stages how sick the liver is and drives transplant priority. It returns the same score regardless of what operation is planned, which is why it discriminated less well for post-operative mortality in the derivation study (30-day C-statistic 0.752 versus 0.859). VOCAL-Penn takes the procedure category and emergency status as inputs. Neither replaces the other: a patient being assessed for surgery usually needs both numbers.
VOCAL-Penn Score vs Child-Turcotte-Pugh
VOCAL-Penn discriminated substantially better for post-operative mortality — Child-Turcotte-Pugh had the weakest performance of every tool compared (C-statistic 0.682 versus 0.859).
Child-Turcotte-Pugh remains a reasonable bedside summary of hepatic function and is embedded in decades of literature, but two of its five inputs (ascites and encephalopathy) are graded subjectively, and it collapses continuous labs into three coarse bands. For a pre-operative risk estimate those properties cost accuracy. Child-Turcotte-Pugh class still matters for drug dosing and for eligibility criteria that are written in terms of it.
VOCAL-Penn Score vs Mayo Risk Score (post-operative mortality in cirrhosis)
VOCAL-Penn outperformed it in the derivation cohort, and the Mayo score's specific failure was calibration — it overestimated risk across the whole spectrum.
The Mayo Risk Score was the previous standard for this question and remains widely cited. In the derivation study its 30-day C-statistic was 0.766 against 0.859 for VOCAL-Penn, and it systematically predicted higher risk than observed, which in practice means patients may have been refused surgery they would have survived. In the smaller external validation cohort the two were closer (0.79 versus 0.82 at 90 days) and the difference was not statistically significant, though VOCAL-Penn still had the better Brier score.
Pearls & pitfalls
- It is not an additive point score — the arithmetic cannot be done at the bedside without a calculator, which is the main reason it is under-used relative to MELD.
- Enter the true ASA class. The authors showed ASA performs better as an ordinal 2/3/4 variable than as a binary one, so collapsing it loses accuracy.
- Platelet count stands in for portal hypertension. Do not also try to account for ascites — the authors found the two near-interchangeable and deliberately chose platelets.
- Laparoscopic versus open abdominal surgery are separate categories, and the difference is large. If both are genuinely options, run the score twice and compare.
- Obesity lowers the predicted risk. That is what the model does, but treat it as an association in this cohort rather than a reason to reassure an obese patient.
- The 90- and 180-day figures are cumulative, not conditional. They already include the earlier risk, so they should never be lower than the 30-day number.
Critical actions
- Do not use the score as the sole determinant of surgical candidacy — it estimates risk, not indication.
- For a patient in the higher predicted range, involve hepatology and anaesthesia before the list date and discuss whether a less invasive approach or pre-operative optimisation changes the estimate.
- Use the 90-day decompensation estimate to plan post-operative monitoring even when the mortality figure looks acceptable — decompensation is the more common outcome.
- Check whether any input falls outside the derivation range; if so, treat the output as an extrapolation and weight it accordingly.
- Consider transplant evaluation rather than elective surgery when predicted mortality is high and the underlying liver disease is progressive.
Why this score exists
The model came out of the Veterans Outcomes and Costs Associated with Liver Disease (VOCAL) group at the University of Pennsylvania, and the design choices are worth knowing because they explain the inputs. The authors tested ascites and found it performed near-interchangeably with platelet count while being far more subjective to assess between clinicians, so platelet count was kept as the marker of portal hypertension and ascites was left out. They also showed ASA class carries more information as an ordinal 2/3/4 variable than collapsed into a binary, which is why the calculator asks for the true class. And they chose to model surgery category explicitly — the gap this fills — after observing that the same patient's risk differs substantially between a laparoscopic and an open abdominal approach. The authors are explicit that the score estimates risk and does not establish whether an operation is indicated.
About the creator
First author, 2021 derivation study
Derived the VOCAL-Penn score for surgical risk in cirrhosis, improving on generic risk models by using cirrhosis-specific predictors.
Senior author
Led the Veterans Affairs cirrhosis outcomes programme from which the cohort came.
Limitations
- Derived and internally validated in a US Veterans Affairs cohort that was 97.2% male; a restricted analysis in women kept discrimination above 0.89, but the female sample was small.
- ASA classes 1, 5 and 6 were excluded, so the score does not apply to the healthiest or the moribund.
- Predictions outside the derivation ranges (age 40–85, albumin 1.5–5 g/dL, bilirubin 0.2–5 mg/dL, platelets 30–450 ×1,000/µL) are extrapolations, and the calculator flags them as such.
- Obesity appearing protective is counterintuitive and may reflect residual confounding rather than a causal benefit.
- The model predicts risk; it does not tell you whether the operation is indicated, nor which alternatives exist.
If you are the patient
If your doctors have used this score, they are trying to answer one specific question: given that you have cirrhosis, how risky is this particular operation for you? The score uses your age, three blood test results, your general fitness for anaesthesia, and — importantly — which operation is planned, because the risk genuinely differs between procedures. It gives a percentage rather than a yes or no, and it does not decide whether you should have surgery. It is a starting point for a conversation with your liver specialist, surgeon and anaesthetist about whether the operation is worth its risk, whether a less invasive version is possible, and what could be improved beforehand. A higher number is not a refusal; it usually means more people need to be involved in the decision.
Frequently asked questions
What is the VOCAL-Penn score?#
The VOCAL-Penn score is a validated risk model that predicts 30-, 90- and 180-day mortality after surgery in patients with cirrhosis, plus the risk of hepatic decompensation within 90 days. It was derived from the Veterans Outcomes and Costs Associated with Liver Disease (VOCAL) cohort at the University of Pennsylvania and published in Hepatology in 2021.
Where can I calculate the VOCAL-Penn score online?#
You can calculate it free on this page — enter age, albumin, bilirubin, platelet count, ASA class, surgery type, emergency status, MASLD status and whether BMI is 30 or above, and the calculator returns 30-, 90- and 180-day mortality along with 90-day decompensation risk. No sign-up is required and the calculation runs entirely in your browser.
What does VOCAL-Penn stand for?#
VOCAL stands for Veterans Outcomes and Costs Associated with Liver Disease, the multicentre cohort the model was built from; Penn refers to the University of Pennsylvania, where the model was developed.
Which variables does the VOCAL-Penn score use?#
Nine: age, serum albumin, total bilirubin, platelet count, ASA physical status class, surgery category, emergency indication, MASLD/NAFLD aetiology, and obesity defined as BMI ≥ 30 kg/m².
How accurate is the VOCAL-Penn score?#
In the internal validation cohort the 30-day C-statistic was 0.859 (95% CI 0.809–0.909), compared with 0.766 for the Mayo Risk Score, 0.752 for MELD-Na, 0.724 for MELD and 0.682 for Child-Turcotte-Pugh. Calibration was also good at each time point, whereas the Mayo Risk Score overestimated risk across the whole spectrum. The model was subsequently externally validated in two large independent health systems.
Is VOCAL-Penn better than MELD or the Mayo Risk Score for surgical risk?#
For predicting post-operative mortality in cirrhosis, yes — it discriminated better than MELD, MELD-Na, Child-Turcotte-Pugh and the Mayo Risk Score at every time point studied, and it is the only one of those tools that accounts for the type of operation. MELD and Child-Turcotte-Pugh remain the right tools for staging liver disease severity in general; they were simply not built to answer a surgical question.
Can the VOCAL-Penn score be used for non-abdominal surgery?#
Yes. The derivation cohort included abdominal, abdominal wall, vascular, major orthopedic and chest/cardiac procedures, and surgery category is an explicit input, so the score applies to both hepatic and non-hepatic operations.
Does the VOCAL-Penn score include ascites or MELD?#
No. The authors tested ascites and found it performed near-interchangeably with platelet count while being more subjective to assess, so platelet count was kept instead. MELD is not an input either — VOCAL-Penn is a standalone model, not a MELD modifier.