We use cookies to ensure our website works properly and to personalise your experience. Cookies policy
Gahlot Institute of pharmacy, Kopar-khairane, Navi Mumbai.
Older adults face a significantly heightened risk of adverse drug reactions (ADRs) and adverse drug events (ADEs) due to physiological changes such as diminished renal and hepatic function, altered pharmacokinetics, and the cumulative burden of multiple chronic diseases. Compounded by widespread polypharmacy, these factors contribute to a disproportionate rate of drug-related complications in geriatric populations. Alarmingly, a substantial proportion of these events—ranging from 30% to 50%—are preventable, underscoring an urgent need for preemptive strategies in prescribing and monitoring.Risk prediction models (RPMs) offer a promising avenue to identify individuals at risk before harm occurs, enhancing clinical decision-making and medication safety. This systematic review and meta-analysis synthesize RPMs specifically developed or validated for predicting ADRs and ADEs in adults aged 60 years and above. Following PRISMA guidelines, we assessed study quality using PROBAST and reporting transparency with TRIPOD.We identified over 30 unique RPMs across various healthcare settings. Most models employed traditional statistical methods such as logistic regression, while a growing subset utilized machine learning algorithms. Common predictors included polypharmacy, impaired renal function, frailty, and prior drug-related incidents. Model performance varied considerably, with area under the curve (AUC) values ranging from 0.60 to 0.85. However, external validation was rare, and calibration metrics were inconsistently reported. Despite their potential, RPMs remain under-integrated into clinical workflows due to limitations in usability, transparency, and generalizability. This review highlights both the promise and shortcomings of current approaches, advocating for the development of robust, explainable, and interoperable models. When appropriately implemented, these tools could mark a pivotal shift toward proactive and personalized pharmacotherapy in older adults.
In recent years, the makeup of healthcare systems around the world has changed a lot(1). Older people, especially those aged 60 and above, are now a major part of the patient group in healthcare(2). These people often face many health problems because of aging, other health conditions, and long-term use of medications(3). While drugs are mostly used to treat chronic diseases in older adults, they can also be a big cause of harm(4). The rise in bad reactions to drugs and harmful drug events has made it important to look closely at how medications are given, watched, and managed in this group(5). Adverse drug reactions are harmful effects that happen even when medicines are taken at the usual dose(6). These can lead to serious health problems, emergency visits, or longer hospital stays(7). Adverse drug events include a range of problems like wrong dosages, not taking medicine as directed, and interactions between drugs or between drugs and diseases(8). About 10 to 30% of hospital visits for older adults are linked to drug problems, and a big part of that is avoidable(9). These numbers show the need for strategies that prevent harm before it happens(10). One main reason for drug problems in the elderly is taking many medications at once, often called polypharmacy(11). Although this is sometimes necessary due to multiple health issues, it can cause conflicts between drugs, make it harder to follow the medication plan, and challenge both the patient’s. ability to take medicine and the doctor’s ability to keep track(12). Aging also brings physical changes like slower kidney function, less liver processing, and changes in body makeup. These changes affect how drugs move through the body and are removed(13). This means medicine needs to be given more carefully and monitored closely(14). Other factors like memory problems, frailty, vision or hearing issues, and lack of support from family or caregivers add more complexity to managing medications in older adults(15). Because standard ways of prescribing aren’t always enough and each person’s risk can be different, doctors and researchers have started using tools called risk prediction models (RPMs)(16). These models use patient data like age, medical history, lab results, and medication lists to guess the chances of future health problems, such as drug reactions or harmful events(17). They can use basic statistical methods like logistic or Cox regression or more complex machine learning methods that can handle different types of data(18). These models can help doctors catch high-risk patients early, check medications, and make better decisions about treatment, especially for older adults. In theory, using RPMs in electronic health records and decision support systems could make prescribing drugs safer and more precise(19). But in practice, this idea is still not widely used. Many models are based on data from just one hospital or a small group, and they may not be clear or well tested. Some models don’t account for other factors, handle missing data poorly, or aren’t well checked for accuracy. Also, even when models are good, doctors may not trust them if they can’t understand how they work. These issues make it hard to use these models and show the need for better methods, clearer standards, and designs that work well with doctors. Besides technical issues, using RPMs in real-world settings faces other challenges(20). Doctors may not know about existing models, may not be trained to use them, or find them hard to fit into their normal work. Support from hospitals, rules from regulators, and input from those involved in healthcare are needed to help bring these models into practice. There is also a growing need for models that are not only accurate but also easy to understand, fair, and can work in different healthcare settings. Testing these models in different populations is important because results from one group may not work well in another group with different health issues or systems. Even though many RPMs have been developed in recent years, the area of geriatric care still lacks a full look at all these models. Most older reviews have looked at how common drug reactions are, the effect of taking many medications, or ways to watch for drug side effects, but few have focused directly on the models themselves—how they are made, what they use, how they are reported, and whether they can be used real-world. Without this, doctors and decision-makers don’t have a clear way to tell which models work well and which don’t. To fill this gap, this review takes a careful look at existing RPMs that predict drug reactions and events in older adults. The main goal is to list these models, check how strong they are using the PROBAST tool, and how clearly they are reported using the TRIPOD guidelines. We also look at how well these models perform using key measures like how good they are at predicting outcomes, how accurate they are, and how well they are tested(21). We look at what variables the models use, the settings in which they are used, and the results they focus on, to find common trends, missing areas, and where more research is needed. This review also gives a chance to think more widely about using predictive models in elderly care, including whether they can be used in different situations, the moral questions they bring up, and how easy they are to fit into existing systems. By showing the strengths and weaknesses of these models, we hope to help doctors, researchers, and system designers improve drug safety for older adults. In doing so, we aim to move the focus from dealing with harm after it happens to preventing it before it occurs, using strong science and caring for people.
This systematic review and meta-analysis was done to look at risk prediction models (RPMs) that help find bad reactions to drugs (ADRs) and bad drug events (ADEs) in older adults(22). The process followed known guidelines to make sure everything was done carefully, clearly, and in a way that others can repeat. The review followed the PRISMA statement, which is a standard for how to report systematic reviews and meta-analyses(23). A thorough plan was made to find all relevant studies published in serious academic journals. Several online databases were checked, like PubMed/MEDLINE, Embase, Scopus, Web of Science, and CINAHL(24). The search didn’t limit the language used and included all articles from the start of these databases up to the latest dates(25). Words and terms related to the topic were combined using “and” and “or” to make the search more accurate(26). These terms included “risk prediction model,” “adverse drug reaction,” “adverse drug event,” “older adults,” “geriatrics,” “polypharmacy,” and “drug safety”(27). Also, the reference lists of the articles found and other reviews were hand-searched to find any studies missed during the electronic search(28). Studies were included if they looked at people 60 years or older and created or tested a model to predict ADRs, ADEs, hospital stays because of drugs, or deaths from medication(29). Both studies that just watch what happens (observational) and those that test treatments (interventional) were considered, as long as they used models with performance information(30). Studies that focused only on how common ADRs or ADEs are, without predicting them, were left out(31). Editorials, opinions, conference summaries, and commentaries were also left out because they didn’t have enough detailed information or validation(32). Two reviewers did the first check of titles and abstracts, then looked at full articles for possible inclusion(33). When they had differences in opinion, they talked it over or brought in a third reviewer(34). A standard form was used to collect data, and it was checked beforehand to make sure it worked well(35). The form collected details like the study’s author, year, country, and setting, information about the people in the study, what kind of model was used, what factors were looked at, how outcomes were defined, what statistical methods were used, and how well the model worked (like AUROC, sensitivity, specificity, calibration, etc.)(36). If information was missing or unclear, the researchers who did the original studies were contacted for more details. The quality of the models and any bias were checked using the PROBAST tool(37). This tool looks at four main areas: the people in the study, the factors used, the outcomes measured, and the analysis done. Each area was rated as having low, high, or unclear risk of bias based on specific rules. If a study had high risk in one or more areas, it was considered possibly biased(38). To check how clear the studies were, the TRIPOD statement was used(39). This looks at how well the models were described, how missing data was handled, and how the models were tested internally and externally. A meta-analysis was done for studies that had enough information, especially those that reported the area under the receiver operating characteristic curve (AUROC)(40). To account for differences between studies, a random-effects model was used(41). The I² statistic was used to check how much the results varied, with values over 50% showing a lot of difference(42). Sensitivity analyses were done to see how study quality, setting (like hospital vs. community), type of model (like statistical vs. machine learning), and how the outcome was defined affected the results(43). Funnel plots were used to check for publication bias, but the interpretation was careful because the studies were observational and varied in how they were reported(44). Calibration, which shows how well a model’s predictions match actual results, wasn’t well reported in many studies. However, whenever possible, measures like the Hosmer–Lemeshow test, calibration slope, and plots comparing predicted and actual risk were considered(45). Since calibration data wasn’t always available, a full analysis wasn’t done, but summaries were included. To keep things consistent and reduce bias, all steps — like data collection, checking the quality of the studies, and combining results — were done by trained reviewers who knew a lot about clinical research and statistics(46). The final data set was checked for accuracy and completeness before analysis(47). All analyses were done using software like R and RevMan. The results were reported following PRISMA guidelines, including a detailed diagram showing the study selection process, tables with model details and performance, and charts that show AUROC results(48). This method was chosen to make sure the findings were strong, repeatable, and useful in real-world settings. By using a structured review and combining it with a statistical analysis, this study gives a clear and reliable look at the RPMs used to predict ADRs and ADEs in older adults.
This systematic review uncovered a wide-ranging collection of studies focused on building and testing risk prediction models (RPMs) aimed at identifying adverse drug reactions (ADRs) and adverse drug events (ADEs) among older adults(49). Following careful screening and eligibility checks, 37 studies met the inclusion criteria. These studies spanned diverse regions—including North America, Europe, Asia, and Oceania—capturing different healthcare systems and prescribing patterns(50).
Most of the models were developed for hospitalized older adults, followed by those in long-term care and community settings. This pattern highlights the heightened vulnerability of older adults during hospital stays and transitions of care. Hospital-based studies benefited from richer clinical data, which likely contributed to stronger predictor variables and more robust model performance.
Across these studies, a total of 32 unique RPMs were identified. These models aimed to estimate the likelihood of experiencing ADRs, ADEs, drug-related hospitalizations, or even mortality due to medication use. While some models addressed general drug-related risks, others were designed around specific drug categories—like anticoagulants, psychotropics, or NSAIDs—or specific at-risk subgroups such as patients with kidney dysfunction, cognitive decline, or frailty. This variation reflects the complex nature of drug risks in aging populations(51).
Common predictors across models included demographic and clinical factors. Age alone was rarely a strong predictor; instead, associated health conditions, lab values, and medication data played a more influential role. Polypharmacy—often defined as taking five or more drugs—stood out as a near-universal risk factor. Some models refined this further by including drug burden scores that considered pharmacological complexity, not just drug count(52).
Models that incorporated lab results often relied on markers of kidney and liver function—such as creatinine clearance, eGFR, or liver enzymes—to estimate drug metabolism and toxicity risk. Frailty, whether diagnosed clinically or inferred through algorithms, consistently correlated with higher ADE risk. Other important variables included cognitive decline, recent hospital admissions, previous ADRs, and living in assisted care facilities(53).
Statistical methods used to build these models varied. Logistic regression was most common, favored for its simplicity and ability to handle binary outcomes. Some studies used Cox regression for time-to-event analyses involving long-term outcomes. Newer studies explored machine learning methods like decision trees, random forests, support vector machines, and neural networks. These advanced methods aimed to detect complex, non-linear patterns, though their real-world use remains limited due to concerns about transparency and interpretability.
Model performance was typically assessed using the area under the receiver operating characteristic curve (AUROC or AUC). Reported AUC values ranged from 0.60 to 0.85. Traditional statistical models achieved moderate accuracy, with AUCs generally between 0.70 and 0.78. Machine learning-based models sometimes showed higher AUCs (above 0.80), but most lacked external validation, limiting confidence in their generalizability(54).
Unfortunately, calibration—the extent to which predicted risks matched actual outcomes—was reported in less than half of the studies. Where reported, calibration was assessed using the Hosmer–Lemeshow test, calibration plots, or slope measures. Better-calibrated models tended to come from larger datasets with strong internal validation, while poorly calibrated models often lacked proper adjustment for confounders or were overfit due to small sample sizes or too many predictors(55).
Validation practices were also inconsistent. Most studies performed internal validation using methods like bootstrapping or cross-validation. External validation, however, was rare. Only seven studies validated their models in new populations, and just three did so in entirely different clinical settings. This lack of external testing presents a major hurdle for real-world application and trust in these models(56).
Study reporting was evaluated using the TRIPOD checklist. While most studies met basic criteria—such as defining predictors and outcomes—many fell short in key areas like sample size justification, handling of missing data, and plans for updating models over time. Only 12 studies described their approach to missing data, and less than half used imputation. These gaps undermine reproducibility and may introduce bias(57).
The risk of bias in the studies was assessed using the PROBAST tool. The highest concern was in the "analysis" domain. Common issues included overfitting, failure to adjust for confounding factors, and unclear variable selection methods. Most studies clearly defined predictor variables, though some included data collected after the outcome began—potentially introducing bias. Definitions of ADRs and ADEs also varied—ranging from formal criteria like the Naranjo Scale to subjective clinical judgments—making comparisons difficult and hampering meta-analyses(58).
A few studies went a step further and assessed real-world implementation or clinical usability. These were mostly pilot efforts that tested how the models could be integrated into electronic health records or explored clinician feedback. Findings emphasized that simple design, interpretability, and ease of use were critical(59). Models that required manual data input or provided vague risk scores were poorly received. In contrast, those integrated into digital systems with automatic alerts were more likely to be adopted.
A meta-analysis was performed using data from 16 studies that reported AUROC values. The pooled AUROC was 0.74 (95% CI: 0.70–0.78), indicating moderate discrimination. Subgroup analyses showed hospital-developed models performed better (AUC: 0.77) than community-based ones (AUC: 0.70), likely due to richer inpatient data. Models that included frailty and kidney function indicators outperformed those relying solely on medication count. There was significant heterogeneity (I² = 65%) among models, driven by differences in study populations, outcome definitions, and statistical approachs(60).
When examining predictors across studies, polypharmacy remained a standout(61). The odds of experiencing an ADR with five or more medications ranged from 2.0 to 3.5, depending on health status and types of drugs(62). Impaired kidney function (eGFR <60 ml/min/1.73m²) was linked to a 1.8–2.4 times higher risk of ADEs, particularly for drugs cleared by the kidneys. Frailty indicators were associated with a 1.6–2.8-fold increase in risk(63). Other consistent predictors included previous ADRs, recent hospital stays, and the use of high-risk medications (as listed in Beers Criteria or STOPP/START tools)(64).
Some studies explored newer predictors like pharmacogenetics, social risk factors, and real-time monitoring data. While intriguing, these were included in only a few models and lacked sufficient evidence for widespread use. Challenges around data standardization, availability, and privacy also limit their practicality for now(65).
A small group of models focused specifically on predicting preventable ADRs—those due to prescribing errors rather than drug toxicity. These models tended to have lower AUCs but were more actionable clinically, offering insights that could prompt deprescribing, dose changes, or medication reviews(66).
In summary, the current field of RPMs for predicting ADRs and ADEs in older adults holds considerable potential but is marked by inconsistencies. Many models are promising in theory, with good predictors and decent performance, yet few are ready for seamless integration into clinical practice. Improving methodological rigor, validation efforts, and transparency in reporting will be essential for these models to transition from research tools to valuable clinical aids(67).
This detailed synthesis not only quantifies how well current models perform but also qualitatively assesses their construction and real-world potential. In the discussion that follows, we will place these findings in the broader context of medication safety in older adults—highlighting how RPMs can evolve into impactful tools for clinical decision-making(115).
DISCUSSION
The findings of this review highlight both the promise and the persistent limitations of existing risk prediction models (RPMs) designed to identify adverse drug reactions (ADRs) and adverse drug events (ADEs) in older adults(69). This is a population that remains disproportionately affected by medication-related harm, owing to a complex interplay of physiological vulnerability, polypharmacy, comorbidity burden, and social factors(70). Prediction models, by virtue of their capacity to synthesize clinical data and generate individualized risk estimates, offer a path forward in addressing this challenge. Yet, as this review demonstrates, the journey from model development to real-world implementation is riddled with methodological, technical, and practical hurdles.
One of the most consistent observations in this synthesis is the central role of polypharmacy as a predictive factor. Nearly every model evaluated incorporated polypharmacy as a variable, and in many cases, it was the strongest or most consistent predictor of ADRs and ADEs(71). This is not surprising, given the robust body of evidence linking the number of medications to increased risk of interactions, cumulative toxicity, and prescribing errors. However, reliance on simple medication counts as a surrogate for complexity has its limitations. Some models attempted to refine this measure by including drug class risk (e.g., anticoagulants, psychotropics), employing scoring tools like the Drug Burden Index, or integrating criteria from instruments like Beers and STOPP(72). These refinements are critical, as they offer a more nuanced view of pharmacotherapy risk than numerical thresholds alone. Nevertheless, the inconsistency in how polypharmacy is defined and operationalized across models remains a barrier to comparability and generalizability.
Renal and hepatic function markers also emerged as important predictors, particularly in models that had access to laboratory data. Renal impairment, often indexed by estimated glomerular filtration rate (eGFR), was strongly associated with increased drug-related harm, especially in patients receiving medications cleared through the kidneys. Similarly, hepatic markers such as ALT, AST, and bilirubin provided insight into metabolic competence, although their predictive power was less consistently demonstrated. Frailty, a multifactorial condition encompassing diminished physiological reserve, appeared frequently as a variable and showed strong associations with ADE risk. However, the methods used to assess frailty varied widely, ranging from clinical scales to indirect proxies like weight loss, falls, and functional impairment. The lack of consensus on frailty measurement may limit the reproducibility of these findings but also underscores the importance of this domain in risk stratification(73).
Machine learning techniques, while less prevalent than traditional statistical approaches, have shown growing appeal in recent years. The few models that incorporated algorithms such as decision trees, random forests, and support vector machines reported higher discriminative performance in terms of area under the curve (AUC). However, these gains were tempered by concerns about transparency, interpretability, and clinical trust. In most cases, the “black box” nature of machine learning models was cited as a potential impediment to clinician acceptance, especially in settings where explainability and accountability are paramount. The debate between accuracy and interpretability is not merely academic—it has real implications for model deployment, particularly when tools are integrated into electronic health records or used for automated alerts. In this context, “explainable AI” emerges as a crucial area of development, aiming to bridge the gap between predictive sophistication and human comprehension(74).
Calibration was an area of notable weakness across models. Even when discrimination (as measured by AUC) was acceptable, many models failed to demonstrate adequate calibration—that is, the agreement between predicted probabilities and observed outcomes. Poor calibration can lead to underestimation or overestimation of risk, potentially undermining clinical decision-making(75). The underreporting of calibration measures, such as the Hosmer–Lemeshow test, calibration slope, or graphical plots, further limits the confidence with which these models can be applied in practice. Calibration is particularly important when models are intended for use across diverse settings, where baseline risks may vary due to population characteristics, prescribing habits, or healthcare infrastructure(76).
External validation was rare, and its absence represents one of the most significant limitations in the current RPM landscape. A model’s performance in the environment in which it was developed offers limited insight into how it will function elsewhere. Without testing in independent populations, the utility of an RPM in broader practice remains speculative. Among the models that underwent external validation, performance often declined compared to internal metrics, reaffirming the importance of this step. The challenge of external validation is partly logistical, requiring access to harmonized datasets and consistent definitions, but it is also conceptual—models must be designed with transportability in mind, not merely tuned to local idiosyncrasies.
Transparency in reporting was another domain where improvement is needed. The TRIPOD checklist provides clear guidance on what constitutes thorough and responsible model reporting, yet many studies fell short in key areas(77). Notably, handling of missing data was poorly documented. In geriatric populations, incomplete information is common due to cognitive impairment, inconsistent documentation, and fragmented care. How models address this issue—whether through imputation, exclusion, or proxy variables—has profound implications for bias and validity. Furthermore, variable selection strategies, performance metrics, and model updating procedures were often underreported, making it difficult to fully assess methodological robustness or to replicate findings.
The heterogeneity across models—in terms of outcome definitions, predictor variables, statistical techniques, and validation approaches—posed a significant barrier to synthesis. Some models focused on preventable ADRs, while others included all drug-related harm regardless of causality. The tools used to define and detect ADRs varied, with some relying on validated scales like the Naranjo algorithm, while others used physician judgment or administrative coding(78). This variability complicates not only comparisons but also the potential application of findings to clinical practice. Standardizing outcome definitions and measurement criteria is essential if RPMs are to be effectively compared, refined, and implemented.
Despite these limitations, the review reveals several promising avenues for future development. Models that incorporate dynamic data—such as changes in laboratory results, medication regimens, or functional status over time—are likely to better reflect real-world risk. Static models, based on a single snapshot in time, may miss important transitions or trajectories that signal impending harm. The integration of time-varying predictors, longitudinal tracking, and feedback loops could enhance sensitivity and relevance, particularly in settings like long-term care or chronic disease management(79).
Another area of opportunity lies in integrating social determinants of health into RPMs. Factors such as socioeconomic status, caregiver support, health literacy, and access to services play a substantial role in medication safety but are rarely included in predictive models(80). Capturing and quantifying these variables is challenging but necessary for comprehensive risk assessment. Models that account for these dimensions may offer more holistic and equitable predictions, especially for marginalized or vulnerable subgroups within the older population.
Clinical implementation remains the ultimate test of a prediction model’s utility. Even the most statistically sound RPMs must contend with practical barriers: integration into workflows, compatibility with health information systems, clinician education, and patient engagement. The few studies that explored usability highlighted the importance of simplicity, clarity, and actionable insights. Clinicians expressed greater trust in tools that provided clear explanations for risk estimates, connected predictions to concrete recommendations (e.g., deprescribing alerts), and required minimal manual input. These preferences align with principles of human-centered design and should inform future model development(81).
The review also calls attention to ethical and policy considerations. Prediction models, particularly those deployed at scale, raise questions about data privacy, consent, and algorithmic fairness. Older adults may be disproportionately affected by biased models if training datasets are skewed or if predictors reflect structural inequities. Ensuring that RPMs are transparent, inclusive, and accountable must be part of their development and deployment roadmap.
In conclusion, the existing landscape of RPMs for ADR and ADE prediction in older adults reflects both meaningful progress and critical gaps. Models have evolved in sophistication and scope, with several demonstrating reasonable discriminative performance and plausible predictors. Yet, the underreporting of calibration, lack of external validation, inconsistent definitions, and barriers to implementation temper enthusiasm. Moving forward, the development of robust, explainable, and clinically integrated RPMs will require interdisciplinary collaboration, methodological rigor, and a commitment to equity and usability. Only then can we harness the full potential of predictive modeling to protect one of the most vulnerable segments of our population from medication-related harm.
CONCLUSION
This systematic review and meta-analysis provide a comprehensive evaluation of risk prediction models (RPMs) developed to anticipate adverse drug reactions (ADRs) and adverse drug events (ADEs) in older adults(82). As medication use becomes increasingly complex in aging populations—driven by multimorbidity, physiological decline, and prolonged therapeutic regimens—the need for proactive, evidence-based strategies to mitigate harm is more pressing than ever(83). RPMs have emerged as promising tools in this context, aiming to shift the paradigm from reactive response to preventative surveillance by identifying individuals at heightened risk before adverse outcomes occur.
The findings of this review underscore the considerable variability in model development, methodological rigor, and clinical utility. While a growing number of RPMs demonstrate acceptable discriminatory power, with area under the curve (AUC) values ranging from 0.60 to 0.85, many fall short in calibration, external validation, and integration readiness(84). The reliance on predictors such as polypharmacy, impaired renal function, frailty, and history of prior ADRs reflects well-established risk factors but also suggests a degree of redundancy and lack of innovation across models. Furthermore, the inconsistent use of standardized definitions and outcome measures hinders comparability and restricts the generalizability of findings(85).
The review also reveals critical methodological limitations that impede model trustworthiness and implementation. Poor reporting of sample size justification, handling of missing data, and validation strategies compromises reproducibility and may lead to inflated performance estimates. The underutilization of TRIPOD and PROBAST frameworks further diminishes the transparency and robustness of many existing RPMs(86). Machine learning approaches offer intriguing possibilities for capturing nonlinear relationships and high-dimensional interactions, yet their lack of interpretability remains a barrier to clinical acceptance and regulatory endorsement(87).
Despite these challenges, several promising models demonstrate the potential to enhance patient safety when integrated into clinical decision support systems (CDSS) and electronic health records (EHRs). The real-world deployment of RPMs, however, will require not only technical refinement but also a commitment to user-centered design, stakeholder collaboration, and iterative validation. Models that produce explainable outputs, facilitate workflow integration, and offer actionable guidance stand the greatest chance of transforming prescribing practices and reducing preventable drug-related harm in older adults(88).
Moving forward, the development of RPMs must embrace multidisciplinary collaboration, leveraging insights from clinical pharmacology, geriatric medicine, informatics, and behavioral science. Future models should aim to incorporate dynamic data streams, longitudinal patient trajectories, and context-aware features that reflect the complexity of aging and care environments. Additionally, the inclusion of social determinants of health may offer a more holistic and equitable approach to risk assessment, especially for populations historically underserved or marginalized by conventional healthcare systems(89).
In sum, while RPMs are not yet universally fit for clinical deployment, they represent an evolving frontier in geriatric pharmacovigilance. The insights from this review offer a roadmap for future innovation, emphasizing the need for methodological rigor, transparent reporting, thoughtful validation, and practical applicability. As the healthcare community continues to prioritize patient safety and personalized care, RPMs may well become indispensable tools in optimizing medication use and minimizing harm in older adults. Their successful integration into practice depends not only on data science and predictive accuracy but on empathy, usability, and a shared commitment to elevating the standard of care for a growing and vulnerable population.
The evolution of risk prediction models (RPMs) for adverse drug reactions (ADRs) and adverse drug events (ADEs) in older adults remains a dynamic and underutilized frontier in clinical pharmacology(90). While current models offer valuable insight into risk stratification, the trajectory forward demands a more innovative, interdisciplinary, and patient-centered approach(91). The future of RPMs must be shaped by methodological sophistication, real-world applicability, and a commitment to equity and transparency(92).
One of the most pressing needs is the development of models grounded in large, diverse, and representative datasets(93). Many existing RPMs are derived from single-center or region-specific cohorts, limiting their external validity and perpetuating bias(94). Multicenter, multinational efforts that incorporate heterogeneity in demographics, healthcare systems, and prescribing patterns will yield models with greater generalizability(95). Additionally, more robust external validation must become the norm rather than the exception. RPMs should be routinely tested in independent populations before being considered for clinical use(96).
The incorporation of dynamic, time-varying data represents another critical frontier(97). Most current models rely on static baseline variables, ignoring the evolving nature of health status, medication regimens, and risk profiles(98). Integrating longitudinal data—such as serial lab results, functional decline, or changes in medication use—could vastly improve predictive accuracy and clinical relevance(99). Real-time models capable of updating risk assessments as new information becomes available could support more responsive and adaptive care(100).
Explainable artificial intelligence (AI) offers a bridge between complexity and clinical trust(101). While machine learning techniques provide impressive performance, their lack of transparency remains a barrier to adoption(102). Future models must prioritize interpretability alongside accuracy, enabling clinicians to understand and act on predictions with confidence(103). Visualization tools, rule-based logic, and transparent feature importance mapping are promising directions that enhance usability without compromising sophistication(104)
Social and behavioral determinants of health must also be considered in next-generation RPMs(105). Factors such as caregiver support, medication literacy, cognitive status, income level, and access to healthcare services play a pivotal role in drug safety but are rarely incorporated into existing models(106). Including these variables would not only improve predictive power but also promote fairness and inclusivity in risk assessment, especially for vulnerable and marginalized populations(107).
Finally, implementation science should guide the translation of RPMs into practice(108). The development of clinically intuitive interfaces, seamless integration into electronic health records, and clear links between risk scores and actionable steps are essential for meaningful uptake(109). Engagement with stakeholders—including clinicians, patients, caregivers, informaticians, and health system leaders—is key to ensuring models meet real-world needs and foster sustainable use(110).
In summary, the future of RPMs in older adults hinges not only on statistical ingenuity but also on thoughtful design, contextual awareness, and ethical stewardship(111). By embracing these principles, the next generation of predictive tools can truly transform geriatric pharmacotherapy—shifting the paradigm toward safer, smarter, and more compassionate medication management(112).
Despite the potential of risk prediction models (RPMs) in improving medication safety for older adults, this review uncovered several important limitations and challenges that must be acknowledged. These constraints affect the reliability, generalizability, and clinical applicability of the existing models, and they should inform both the interpretation of current findings and the design of future research(113).
First, there is marked heterogeneity in model design, predictor selection, outcome definitions, and performance metrics across studies. This variability limits the comparability of RPMs and makes meta-analytic synthesis difficult. For example, while polypharmacy is a consistent predictor, its operational definition varies significantly — ranging from simple counts to weighted indices — thus affecting both the strength and direction of its predictive power. Similarly, definitions of ADRs and ADEs differ widely, with some models relying on clinical judgment, others using structured algorithms like the Naranjo scale, and some deriving outcomes from administrative data, all of which pose unique methodological implications(114).
Second, many models suffer from limited external validity. The vast majority are developed in single-center cohorts or narrowly defined populations, which restricts their generalizability. External validation — a critical step in model appraisal — is inconsistently reported or entirely absent in most studies. Without independent validation, it is unclear whether a model's performance will hold in new settings, across diverse patient populations, or within different healthcare systems(115).
Third, transparency in reporting remains a major concern. Several studies do not adhere to standardized guidelines such as TRIPOD, resulting in incomplete information on data handling, statistical methodology, and model implementation procedures. Details about missing data strategies, variable selection techniques, overfitting prevention, and calibration are either underreported or omitted, undermining reproducibility and interpretability. This lack of transparency hinders clinician trust and impedes regulatory review and adoption(116).
Another noteworthy challenge is the underrepresentation of time-varying data and dynamic predictors. Most RPMs are based on static snapshots of patient health, failing to account for evolving clinical trajectories such as changes in medication exposure, laboratory values, or functional status. Without temporal granularity, these models risk missing early warning signs or transient risk spikes, particularly in patients with complex, fluctuating conditions(117).
The limited use of social and behavioral determinants also constrains model comprehensiveness. Medication adherence, cognitive status, caregiver involvement, and socioeconomic factors profoundly influence ADR/ADE risk, yet are rarely integrated into predictive frameworks. This omission may reduce accuracy and contribute to inequities, particularly among vulnerable populations(118).
Finally, implementation challenges persist. Even well-designed models may face barriers in clinical settings due to workflow disruptions, alert fatigue, resistance to change, and technological limitations. RPMs that are not seamlessly embedded into electronic health records or fail to provide actionable guidance may languish unused, regardless of their theoretical utility(119).
In sum, while RPMs offer a valuable lens for anticipatory drug safety in older adults, their current limitations warrant caution. Continued efforts are needed to enhance methodological rigor, validation practices, transparency, and usability to ensure that these tools fulfill their intended promise in real-world care.
REFERENCES
Anjali Shirse, Saee Sutar, Dr. Sandeep waghulde, Risk Prediction Models for Adverse Drug Reactions and Adverse Drug Events in Older Adults: A Systematic Review and Meta-Analysis, Int. J. of Pharm. Sci., 2026, Vol 4, Issue 4, 2145-2163, https://doi.org/10.5281/zenodo.19565017
10.5281/zenodo.19565017