23:40:54 – Development and Internal Validation of a Multimodal Prediction Model for Moderate-to-Severe Osteoporotic Vertebral Compression Fractures Using Bone Quality Biomarkers Derived From CT and MRI

Rationale and Objectives

To develop and internally validate interpretable prediction models using pre-fracture CT-derived Hounsfield unit (HU) values and MRI-based vertebral bone quality (VBQ) scores to estimate the risk of moderate-to-severe osteoporotic vertebral compression fractures (OVCFs).

Materials and Methods

This retrospective study included 256 patients with first-onset, single-level OVCFs from 2022 to 2025. Patients were randomly allocated to training ( n = 180) and testing ( n = 76) sets at a 7:3 ratio. Moderate-to-severe compression was defined as Genant grade ≥ 2. Four prespecified logistic regression models assessed the incremental value of VBQ and HU beyond baseline clinical predictors. Discrimination, calibration, decision curve analysis, threshold behavior, and clinical impact were evaluated. Supplementary machine learning models were benchmarked using the same full predictor set, and the final model was translated into a nomogram.

Results

In testing set, model performance improved progressively with added imaging biomarkers. The full multimodal model achieved the highest discrimination, with an AUC of 0.759 and AUPRC of 0.798, outperforming the clinical baseline model (AUC, 0.587; DeLong P = 0.011) and the base + VBQ model (AUC, 0.630; P = 0.014). Calibration assessment showed a Brier score of 0.205 and a non-significant Hosmer–Lemeshow test (P = 0.679). Decision curve analysis showed greater net benefit within clinically relevant threshold ranges. Supplementary machine learning algorithms did not outperform the logistic regression framework.

Conclusion

A multimodal model incorporating HU and VBQ achieved the best overall performance and was translated into a nomogram to support individualized risk estimation for moderate-to-severe OVCFs.

Take Home Message

Pre-fracture CT-derived HU and MRI-based VBQ improved risk stratification for moderate-to-severe OVCFs beyond clinical predictors, with HU providing the dominant incremental signal.

The full HU–VBQ model showed the best overall performance, but VBQ’s added value beyond HU was not statistically confirmed and requires external validation.

INTRODUCTION

Osteoporotic vertebral compression fractures (OVCFs) are common fragility fractures in older adults and impose a substantial healthcare burden worldwide , . Moderate-to-severe OVCFs (Genant grade ≥2) are associated with severe back pain, progressive kyphosis, functional decline, and poorer outcomes , . Identifying patients at risk of severe collapse may support earlier intervention, closer monitoring, and individualized treatment planning.

Current assessment of fracture severity is usually post-fracture and morphology-based. Although useful for characterizing established deformity, these parameters cannot inform pre-fracture risk stratification , . Imaging-derived bone quality markers may therefore provide a biologically relevant assessment of vertebral vulnerability.

Because vertebral strength is determined not only by mineral density but also by trabecular microarchitecture, reliance on computed tomography (CT)-based attenuation alone may be insufficient. The magnetic resonance imaging (MRI)-based vertebral bone quality (VBQ) score has recently emerged as a promising biomarker that reflects trabecular microstructural integrity and marrow composition ,, . Hounsfield unit (HU) values derived from routine CT have been used as an opportunistic measure of local bone mineral density, and prior studies have suggested an association between lower HU values and more severe vertebral compression ,, . Whereas HU primarily reflects bone mineral density, VBQ captures information on trabecular microstructure and marrow composition, suggesting that the two modalities may provide complementary insights into vertebral fragility.

Most previous investigations have focused on single-modality imaging or simple association analyses . However, the combined predictive contribution of CT-derived HU and MRI-based VBQ to the severity of subsequent OVCFs has not been fully clarified. In the present study, we therefore aimed to develop and internally validate a set of prespecified clinical prediction models to quantify the incremental value of HU and VBQ beyond a clinical baseline model.

Using pre-fracture assessments obtained within 12 months before the index fracture, we evaluated routinely available clinical factors and bone quality biomarkers while excluding post-fracture morphology. Four prespecified logistic regression models (M1–M4) assessed the added value of HU, VBQ, and their combination. Supplementary machine learning analyses using the same predictor set examined robustness beyond logistic regression. If validated, this approach could support pre-fracture risk stratification.

MATERIALS AND METHODS

Study Design and Patient Cohort

Clinical and imaging data were retrospectively collected from patients with single-level OVCFs admitted to the First Affiliated Hospital with Nanjing Medical University between 2022 and 2025. Eligible patients had acute or subacute painful vertebral compression fractures confirmed by spine X-ray, with bone marrow edema on T2-weighted fat-suppression MRI or increased activity on bone scan. Of 967 initially eligible patients, 711 were excluded because of pathological, infectious, high-energy, or non-osteoporotic metabolic fractures ( n = 165), multilevel fractures ( n = 108), prior vertebral fracture or spine surgery ( n = 202), or missing outcome or imaging data ( n = 236). The final 256 patients were randomly assigned to training ( n = 180) and testing ( n = 76) sets.

Baseline predictors were defined a priori as the most recent complete pre-fracture assessment within 12 months before the index fracture, defined as the date of first radiographic confirmation of the target OVCF. Only pre-index data were eligible, and when multiple visits were available, the closest assessment was selected.

This study protocol complied with the Declaration of Helsinki and was approved by the Ethics Committee of the First Affiliated Hospital with Nanjing Medical University (Approval No. 2025-SR-847) . Because this was a retrospective study and all patient data were strictly anonymized, the requirement for informed consent was waived by the committee.

Outcome Definition

The primary outcome was incident OVCF severity at the index fracture, defined using the Genant semiquantitative system . Moderate-to-severe compression was defined as grade 2 (25–40% height loss) or grade 3 (>40% height loss); the outcome was dichotomized as mild vs. moderate-to-severe compression.

Candidate Predictors and Image Measurements

Variables were obtained from electronic medical records and picture archiving and communication systems, including demographics, comorbidities, lifestyle, and bone quality parameters. Candidate predictors were selected a priori based on prior evidence, biological plausibility, and routine availability, not univariable significance . Age, sex, and body mass index (BMI) were core covariates; hypertension, type 2 diabetes mellitus (T2DM), and steroid history were included for reported associations with bone fragility or fracture severity ,, . Multicollinearity was assessed using the variance inflation factor, with values >5 indicating significant multicollinearity .

Two imaging indicators were included as follows: CT-derived HU and MRI-derived VBQ . HU was measured on clinical CT using a predefined region-of-interest protocol and averaged across three cancellous bone regions. Because imaging was routine and retrospective, no phantom calibration or scanner-specific correction was performed. VBQ was calculated as the median vertebral body signal intensity divided by cerebrospinal fluid signal intensity. Both parameters were treated as pre-fracture surrogates of vertebral bone quality.

All imaging datasets were independently reviewed by two senior spine surgeons blinded to clinical predictors. Agreement for Genant grading was assessed using weighted kappa with 95% confidence intervals (CIs) , and reliability for HU and VBQ was assessed using single-measure absolute-agreement intraclass correlation coefficients (ICCs) from a two-way random-effects model , . Disagreements were adjudicated by a third blinded senior spine surgeon, and final values were determined by consensus.

Data Preprocessing

All preprocessing was performed using the training data only, and derived rules were applied to the testing set . Missingness was assessed before modeling. BMI was the only primary-model predictor with missing values; all other predictors were complete. Missing BMI values were imputed using multiple imputation by chained equations within the training set , with five imputations considered sufficient because missingness was low and limited to one predictor. Estimates were pooled using Rubin’s rules , and the training-derived imputation model was applied to the testing set.

For the primary logistic regression models, no additional feature transformation was required beyond imputation. For the supplementary machine learning analyses, continuous variables were further standardized using z-score normalization, and categorical variables were encoded as binary indicator variables where appropriate.

Primary Model Specification and Development

The primary analysis used four prespecified nested logistic regression models to quantify the incremental value of HU and VBQ beyond the clinical baseline model. M1 included age, sex, BMI, hypertension, T2DM, and steroid history; M2 added VBQ; M3 added HU; and M4 added both HU and VBQ. All models were fitted to the training set. Supplementary machine learning analyses were restricted to the M4 predictor set. With 180 training patients and 8 M4 predictors, the events per variable exceeded 10, satisfying standard recommendations .

Threshold Selection Strategy

The primary analysis focused on comparative performance across the four prespecified nested models rather than on algorithm selection. The testing set remained entirely untouched during model development and preprocessing-rule derivation and was used solely for independent performance assessment.

Thresholds were determined without testing-set leakage using nested cross-validation within the training set. An outer stratified 10-fold loop generated out-of-fold (OOF) predictions, and an inner stratified 5-fold loop repeated preprocessing and model development within each outer training partition. Aggregated OOF predictions were used to construct the receiver operating characteristic (ROC) curve and identify the Youden-optimal threshold, which was then fixed and applied unchanged to the testing set . The testing set was not used for fold assignment, preprocessing, threshold optimization, or model selection.

Model Performance Evaluation

Discrimination was evaluated using ROC curve analysis and precision–recall curve (PRC) analysis. The area under the ROC curve (AUC) and the area under the precision–recall curve (AUPRC), together with their bias-corrected 95% CIs, were reported using 1000 bootstrap resamples . Sensitivity, specificity, accuracy, and F1 score were additionally calculated using the OOF-derived optimal cutoff determined in the training set and then applied unchanged to the testing set.

Calibration was assessed in the testing set using loess-smoothed calibration plots comparing predicted probabilities with observed proportions. Quantitative calibration was further evaluated using the Brier score, calibration slope, calibration intercept, and Hosmer–Lemeshow test with the grouping parameter set to g = 5 .

In the testing set, discrimination across the four nested models (M1–M4) was compared using DeLong’s test . Specifically, M2 vs. M1 was used to assess the incremental contribution of VBQ, M3 vs. M1 was used to assess the incremental contribution of HU, and M4 vs. M1 was used to assess the combined incremental contribution of both imaging biomarkers. P values for pairwise comparisons were adjusted using the Bonferroni method where applicable.

Decision curve analysis (DCA) and clinical impact curves (CIC) evaluated clinical utility across threshold probabilities , . A nomogram based on the final M4 logistic regression model provided individualized prediction. Predictor points were derived from regression coefficients, summed, and mapped to predicted probability using the model intercept and logistic transformation. Nomogram-derived probabilities were computationally checked against model-derived probabilities .

Supplementary Machine Learning Analyses

As a secondary robustness analysis, six machine learning (ML) classifiers were developed using the M4 predictor set: least absolute shrinkage and selection operator (LASSO) regression, elastic net, linear support vector machine (SVM), random forest, extremely randomized trees (ExtraTrees), and extreme gradient boosting (XGBoost). These analyses assessed algorithmic robustness rather than replacing the prespecified logistic regression framework. Models were developed in R 4.2.3 using caret 7.0.1; hyperparameters were optimized within the training set using repeated cross-validation and grid search, and final models were refitted on the full training set.

The supplementary ML analyses used the preprocessed training and testing sets described previously. To further characterize the robustness and predictive behavior of the M4-based supplementary ML models, the following analyses were performed and are presented in the Supplementary Figures : (1) stability of model performance across repeated 10-fold cross-validation; (2) threshold optimization based on training-set OOF predicted probabilities; (3) confusion matrix analysis in the testing set; (4) Kolmogorov–Smirnov (KS) curve analysis; (5) correlation analysis between HU and VBQ in the training set; (6) feature importance ranking; and (7) Shapley additive explanations (SHAP) summary (beeswarm) plots in the training set.

Statistical Analysis

Baseline characteristics were compared in the overall study cohort between patients with mild compression and those with moderate-to-severe compression. Continuous variables with a normal distribution are presented as mean ± standard deviation and were compared using the t-test, whereas non-normally distributed continuous variables are presented as median (interquartile range) and were compared using the Mann–Whitney U test. Categorical variables are presented as numbers (percentages) and were compared using the chi-square test or, when the expected cell count in any category was less than 5, Fisher’s exact test.

To explore the unadjusted and adjusted associations between candidate predictors and moderate-to-severe compression, univariable and multivariable logistic regression analyses were performed in the training set. Results are reported as odds ratios with 95% CIs. These analyses were conducted for descriptive purposes only and were not used for predictor selection or model specification in the prespecified nested models (M1–M4). All statistical tests were two-sided, and a P value <0.05 was considered statistically significant.

RESULTS

Cohort Assembly, Reproducibility, and Dataset Comparability

Of 967 patients with OVCFs screened, 256 met the predefined eligibility criteria and were included in the final analytic cohort ( Fig 1 ). The cohort was randomly divided into a training set ( n = 180) and a testing set ( n = 76) for model development and internal validation. Moderate-to-severe OVCFs (Genant grade ≥ 2) accounted for 51.95% (133/256) of the overall cohort, with similar prevalence in the training and testing sets (51.67% vs. 52.63%), confirming a balanced outcome distribution after random splitting and the suitability of the testing set for internal validation.

Figure 1

Flowchart of patient selection and cohort allocation: a total of 967 consecutive patients with vertebral compression fractures were screened. After predefined exclusions, 256 patients were included and randomly allocated into the training set ( n = 180) and testing set ( n = 76) at a 7:3 ratio.

Measurement reproducibility was confirmed before formal model development. Interrater agreement was good for Genant grading (weighted kappa = 0.83), and interrater reliability was high for both HU (ICC = 0.92) and VBQ (ICC = 0.88) ( Table S1 ), indicating that both the outcome definition and the two key imaging biomarkers were sufficiently robust for subsequent analyses.

Baseline comparability between the training and testing sets was generally preserved ( Table 1 ). The two sets were similar with respect to age, sex, BMI, hypertension, T2DM, history of steroid use, anti-osteoporosis treatment, smoking, alcohol use, HU, and VBQ. Among index fracture descriptors, only Genant deformity subtype differed between groups, whereas vertebral level, Cobb angle, paraspinal muscle fatty infiltration, and outcome prevalence remained comparable. Missingness was limited: height, weight, and BMI had missing values in 24 patients (9.4%; available n = 232), whereas all other variables presented in Table 1 were complete. Among predictors included in the primary models, BMI was the only variable with missing values and was addressed using multiple imputation. Collectively, these findings support the suitability of the testing set for internal validation rather than representing a systematically easier or harder subset.

Table 1

Comparison of Pre-fracture Baseline Variables and Index-Fracture Characteristics Between the Training and Testing Sets

Variable Total ( n = 256) Grade < 2 ( n = 123) Grade ≥ 2 ( n = 133) P value Training set ( n = 180) Testing set ( n = 76) P value
Pre-fracture Baseline Variables
Demographics
Age, years 73.78±8.39 73.13±8.45 74.38±8.32 0.236 73.33±8.50 74.83±8.09 0.185
Sex, n % 0.160 0.077
Female 204 (79.69%) 93 (75.61%) 111 (83.46%) 137 (76.11%) 66 (86.84%)
Male 52 (20.31%) 30 (24.39%) 22 (16.54%) 43 (23.89%) 10 (13.16%)
Height, cm 160.00 (156.00, 165.00) 160.00 (157.50, 167.00) 159.00 (155.00, 164.00) 0.034 160.00 (157.00, 165.00) 159.50 (155.75, 163.25) 0.382
Weight, kg 60.00 (55.00, 67.63) 60.00 (55.00, 69.50) 60.00 (55.00, 65.00) 0.131 60.00 (55.00, 66.25) 62.00 (54.75, 70.00) 0.448
BMI, kg/m 2 23.42 (21.48, 25.49) 23.42 (21.41, 25.43) 23.42 (21.48, 25.54) 0.748 23.23 (21.48, 25.36) 23.63 (21.43, 25.78) 0.328
Comorbidities and Lifestyle
Hypertension, n % 122 (47.66%) 53 (43.09%) 69 (51.88%) 0.200 87 (48.33%) 35 (46.05%) 0.844
T2DM, n % 51 (19.92%) 22 (17.89%) 29 (21.80%) 0.530 35 (19.44%) 16 (21.05%) 0.902
History of steroid use, n % 7 (2.73%) 2 (1.63%) 5 (3.76%) 0.449 6 (3.33%) 1 (1.32%) 0.678
Anti-osteoporosis treatment, n % 2 (0.78%) 0 (0.00%) 2 (1.50%) 0.499 2 (1.11%) 0 (0.00%) 1.000
Current smoker, n % 12 (4.69%) 7 (5.69%) 5 (3.76%) 0.664 11 (6.11%) 1 (1.32%) 0.116
Alcohol use, n % 7 (2.73%) 4 (3.25%) 3 (2.26%) 0.714 7 (3.89%) 0 (0.00%) 0.108
Imaging Parameters
Paraspinal muscle fatty infiltration, n % 0.063 0.269
Light 143 (55.86%) 78 (63.41%) 65 (48.87%) 105 (58.33%) 38 (50.00%)
Moderate 84 (32.81%) 33 (26.83%) 51 (38.35%) 58 (32.22%) 26 (34.21%)
Heavy 29 (11.33%) 12 (9.76%) 17 (12.78%) 17 (9.44%) 12 (15.79%)
HU value 70.00 (50.94, 89.67) 78.00 (57.83, 104.17) 63.00 (46.33, 80.67) <0.001 70.67 (52.69, 89.75) 67.48 (49.06, 87.12) 0.634
VBQ score 3.30 (2.77, 3.80) 3.07 (2.65, 3.68) 3.37 (2.95, 3.83) 0.006 3.33 (2.85, 3.78) 3.18 (2.64, 3.84) 0.311
Index Fracture Characteristics
Fracture Characteristics
Duration, days 6.00 (3.00, 11.00) 5.00 (3.00, 10.00) 7.00 (3.00, 14.00) 0.109 6.00 (3.00, 11.00) 6.00 (3.75, 13.25) 0.323
Location of level 0.749 1.000
Thoracic vertebra, n % 53 (20.70%) 27 (21.95%) 26 (19.55%) 37 (20.56%) 16 (21.05%)
Lumbar vertebra, n % 203 (79.30%) 96 (78.05%) 107 (80.45%) 143 (79.44%) 60 (78.95%)
T-L junction, n % 210 (82.03%) 101 (82.11%) 109 (81.95%) 1.000 144 (80.00%) 66 (86.84%) 0.261
Genant deformity type, n % 0.194 0.042
Wedge 99 (38.67%) 42 (34.15%) 57 (42.86%) 63 (35.00%) 36 (47.37%)
Biconcave 156 (60.94%) 80 (65.04%) 76 (57.14%) 117 (65.00%) 39 (51.32%)
Crush 1 (0.39%) 1 (0.81%) 0 (0.00%) 0 (0.0%) 1 (1.32%)
AO Spine <0.001 0.484
A1 213 (83.20%) 118 (95.93%) 95 (71.43%) 147 (81.67%) 66 (86.84%)
A2 22 (8.59%) 2 (1.63%) 20 (15.04%) 15 (8.33%) 7 (9.21%)
A3 20 (7.81%) 3 (2.44%) 17 (12.78%) 17 (9.44%) 3 (3.95%)
A4 1 (0.39%) 0 (0.00%) 1 (0.75%) 1 (0.56%) 0 (0.00%)
Cobb angle, ° 12.70 (7.75, 20.23) 11.38 (6.84, 17.09) 14.06 (8.78, 21.98) 0.002 12.81 (7.80, 20.43) 12.24 (7.61, 19.16) 0.564
Clinical Outcome
Moderate-to-Severe compression, n % 133 (51.95%) 0 (0.00%) 133 (51.95%) <0.001 93 (51.67%) 40 (52.63%) 0.997
Only gold members can continue reading. Log In or Register to continue
Sep 2, 2026 | Posted by in ANESTHESIA | Comments Off on 23:40:54 – Development and Internal Validation of a Multimodal Prediction Model for Moderate-to-Severe Osteoporotic Vertebral Compression Fractures Using Bone Quality Biomarkers Derived From CT and MRI

Full access? Get Clinical Tree

Get Clinical Tree app for offline access