U.S. flag

An official website of the United States government

NCBI Bookshelf. A service of the National Library of Medicine, National Institutes of Health.

Busse R, Klazinga N, Panteli D, et al., editors. Improving healthcare quality in Europe: Characteristics, effectiveness and implementation of different strategies [Internet]. Copenhagen (Denmark): European Observatory on Health Systems and Policies; 2019. (Health Policy Series, No. 53.)

Cover of Improving healthcare quality in Europe

Improving healthcare quality in Europe: Characteristics, effectiveness and implementation of different strategies [Internet].

Show details

14Pay for Quality: using financial incentives to improve quality of care

, , and .

Summary

What are the characteristics of the strategy?

The main attribute of Pay for Quality (P4Q) is that a financial incentive is paid to a provider or professional for achieving a quality-related target within a specific time-frame. P4Q can be implemented in various healthcare settings, targeting a range of healthcare providers or professionals. P4Q schemes can reward high quality measured in terms of structures, processes and/or outcomes, and/or penalize low quality. P4Q schemes can be implemented in line with other quality improvement interventions.

What is being done in European countries?

The implementation of P4Q schemes began in the late 1990s. A total of 14 primary care P4Q programmes and 13 hospital P4Q programmes were identified in a total of 16 European countries. P4Q schemes in primary care incentivize mostly process and structural quality with respect to prevention and chronic care. P4Q schemes in hospital care incentivize more often improvements in health outcomes and patient safety. The size of financial incentives varies between 0.1% and 30% of total provider income in primary care (individual physicians or primary care practices) and between 0.5% and 10% of total provider income in hospital care.

What do we know about the effectiveness and cost-effectiveness of the strategy?

Overall, the effectiveness and cost-effectiveness of P4Q schemes remains unclear. The most reliable studies of P4Q in primary care suggest small positive effects on process-of-care (POC) indicators, while in-hospital care schemes appear to be ineffective with respect to POC measures. For both settings the evidence on effectiveness with respect to improving health outcomes and patient safety indicators is inconclusive. Patient experience and patient satisfaction were rarely evaluated and if they were, they usually did not improve. In fact, in some primary care programmes chronically ill patients experienced worsened continuity of care. Furthermore, a few studies suggest that P4Q schemes are less effective than other quality improvement initiatives, such as public reporting or audit and feedback.

How can the strategy be implemented?

P4Q schemes are more effective when the focus of a scheme is on areas of quality where change is needed and if the scheme embraces a more comprehensive approach, covering many different areas of care. Quality measures should be developed in collaboration with relevant healthcare professionals and reinforce professional norms and beliefs. Payment mechanisms have to be codified very clearly, with statements of entitlements, conditions, time horizons and criteria for receipt of funds.

Conclusions for policy-makers

While reliable evidence on the effectiveness of P4Q programmes is scarce, there is a broad consensus that such programmes are technically and politically difficult to implement. All relevant stakeholders should be involved in the process of scheme development. The contents and structure of the scheme have to be kept under review and regularly updated, and adverse behavioural responses need to be monitored. More evidence is needed on the comparative effectiveness of P4Q schemes in comparison to other quality improvement initiatives.

14.1. Introduction: the characteristics of pay for quality

Pay for quality (P4Q) initiatives are increasingly used in healthcare systems in Europe and beyond. Interest in P4Q by researchers and policy-makers has seen an incredible growth since the late 1990s, when the first programmes started to emerge in Europe and the USA (Cashin, 2014). However, despite the growth in P4Q programmes, P4Q remains highly controversial for a wide range of conceptual, practical and ethical reasons (Roland & Dudley, 2015; Wharam et al., 2009). In fact, there is no universally accepted definition of P4Q, and the term is often used interchangeably with “pay for performance” (P4P). Yet the term P4Q is more precise, as it makes clear that payment depends on the quality of care – and not on other dimensions of health system performance (see also Chapter 1).

The two characteristic features of P4Q programmes are that (1) performance of providers is monitored in relation to pre-specified quality indicators and (2) a monetary transfer is made conditional on the (achievement or improvement of) measured quality of care. In theory, as discussed in Chapter 3, quality can be measured by use of structure, process or outcome indicators of quality – and this is true also for P4Q programmes. In addition, P4Q programmes can, in theory, aim at assuring or improving quality in different areas of care (preventive, acute, chronic or long-term care), and target different types of professional (for example, physicians, nurses or social workers) and providers (for example, primary care practices, hospital departments or hospitals). Furthermore, quality may be incentivized with the aim of assuring or improving quality in terms of effectiveness, safety and/or responsiveness. Nevertheless, despite the potentially very large variation of different characteristics of P4Q programmes, this chapter shows that most existing programmes target a more narrow set of providers (namely primary care providers and hospitals), and that certain characteristics are much more common in P4Q programmes in primary care than in P4Q programmes in hospital care.

P4Q can be implemented together with other quality improvement strategies, such as audit and feedback (see Chapter 10) and public reporting (see Chapter 13). In fact, by design, a P4Q programme includes elements of audit and reporting, since the performance has to be monitored and performance data have to be transmitted to the programme administrators.

The chapter follows the standard structure of chapters in Part 2 of this book. The next section explains why P4Q is expected to contribute to healthcare quality. The following section provides an overview of a selection of existing national and regional P4Q programmes in Europe based on a rapid review (see Box 14.1 for a summary of the methods). The next section summarizes the available evidence on the effectiveness and cost-effectiveness of existing P4Q programmes in Europe and other high-income countries based on a review of reviews, followed by a discussion of the organizational and institutional requirements for the implementation of P4Q programmes, before we draw together the conclusions of the chapter for policy-makers.

Box Icon

Box 14.1

Review methods used to inform the content of this chapter.

14.2. Why should pay for quality contribute to healthcare quality?

The incentives of provider payment systems are known to have a profound impact on the volume and quality of care (Busse & Blümel, 2015; Conrad & Christianson, 2004; Dudley et al., 1998). However, under traditional payment mechanisms, the incentives for the provision of high or better quality of care are indirect and often incidental. For example, fee-for-service payment creates incentives for high levels of provision, and thus might indirectly lead to higher levels of quality. However, fee-for-service may also lead to overprovision of unnecessary, inappropriate and potentially unsafe services, and potentially may pose a barrier to quality improvement if this leads to lower numbers of services being delivered. In contrast, capitation payments eliminate incentives for overprovision and facilitate expenditure control. However, they do not create incentives for quality – and may even be a barrier for quality improvement – because providers have incentives to skimp on necessary services in order to achieve lower costs. Similar problems arise with two common payment methods in the hospital sector – global budgets and Diagnosis Related Group (DRG)-based case payments (Busse & Blümel, 2015) – as neither provides incentives for quality and instead may even pose barriers to quality improvement.

In this context, the idea of P4Q programmes is to change the incentives for providers (professionals and organizations) and to explicitly reward the provision of high or better quality of care – or to penalize poor quality. The assumption is that providers (professionals and/or organizations) will improve the quality of care – through whatever mechanism – if they have a direct financial interest to do so. However, this assumption is highly controversial (Kronick, Casalino & Bindman, 2015). Proponents of P4Q (and P4P more generally) believe that quality improvement strategies relying exclusively on intrinsic motivation of providers (for example, audit and feedback; see Chapter 10) or on non-financial incentives (for example public reporting; see Chapter 13) are insufficient to motivate quality improvements (Rosenthal et al., 2004). Opponents believe that financial incentives could crowd out the intrinsic motivation of physicians to provide high-quality care and could potentially have adverse consequences, such as an exclusive focus on incentivized quality measures while disregarding other potentially important areas of quality (Kronick, Casalino & Bindman, 2015).

The theory underlying many P4Q programmes can be traced to the economic principal/agent literature (Christianson, Knutson & Mazze, 2006; Conrad, 2015; Robinson, 2001). According to the theory, a principal (usually a strategic purchaser) wishes to structure the contractual relationship with the agent (either an individual practitioner or an organization) to secure high-quality health services. It is assumed that increasing quality requires “effort” on the part of the agent, who must therefore be compensated with a financial reward if improvements are to be secured. The agent will then assess how much effort to exert by comparing the expected financial benefits to the effort required. In the simplest form of this model, the principal then sets the financial rewards for the agent knowing how the agent will respond to the incentives, in terms of exerting increased effort, and thereby delivering improved quality. In setting the incentive regime, the principal must of course balance the expected costs of the rewards against the expected improvements in quality.

As set out by Cashin (2014) there are several elements in this model that require more detailed scrutiny. First, measurement plays a key role. Effort cannot usually be observed and measured, so instead there must be some way of explicitly measuring the quality attained. Quality indicators therefore play a key role in any P4Q programme. Ideally these should be accurate and timely indicators of the desired quality criterion, sensitive to variations in provider effort, and resistant to manipulation or fraud. In examining the programmes described in this chapter, it is important to assess the strengths and limitations of the quality metrics being used (see also below).

Second, design of the financial reward mechanism requires numerous judgements, such as the magnitude of the rewards, how they increase with increased quality, whether or not the rewards are based on performance relative to other providers, whether rewards are based on individual aspects of performance or on an aggregate measure of organizational attainment, and whether they are based on absolute levels of attainment or on improvements from previous levels (Eijkenaar, 2013). These design considerations are a central concern of all P4Q programmes, and are likely to play a crucial role in their effectiveness. They are described in Box 14.2 and discussed in more detail later in this chapter.

Box Icon

Box 14.2

Structures of financial incentives within P4Q.

Third, the effect of any P4Q scheme depends crucially on the intrinsic motivation of the professionals and organizations at whom the programme is directed. If the desired improvements in quality are aligned with professional objectives, and the programme serves to offer focus and encouragement to professionals and organizations seeking to secure such improvements, then it may indeed contribute to the desired outcomes. However, if the P4Q programme contradicts or undermines professional motivation, it may prove ineffective or even lead to adverse outcomes.

More generally, it is likely that contextual factors play a key role in the success or otherwise of P4Q programmes. Some aspects of health services are more amenable to P4Q than others, for example those for which reliable performance metrics can be developed. Furthermore, professionals and provider organizations may require a long-term commitment from payers to the P4Q before they are prepared to commit resources to quality improvement efforts. Finally, a persistent theme found throughout the P4Q literature (for example, Damberg et al., 2014; Kane et al., 2004; Kondo et al., 2016; Milstein & Schreyoegg, 2016; Scott et al., 2011) is that effective governance arrangements are an essential prerequisite for the success of any scheme. These have to ensure that information is reliable, that providers are not “cherry-picking” patients who are expected to secure high-quality outcomes, and that non-incentivized aspects of care remain satisfactory.

14.3. What is being done in Europe?

Our review (see Box 14.1) identified a total of 27 P4Q programmes that have been implemented in 16 European countries in both primary and hospital care (Tables 14.1 and 14.2). To our knowledge, the first nationwide P4Q programme introduced in Europe was the Incitant Qualité, implemented in Luxembourg in 1998 (FHL, 2012). We did not identify any P4Q programmes focusing on palliative care.

Table 14.1. Identified P4Q programmes in primary care in eleven EU countries.

Table 14.1

Identified P4Q programmes in primary care in eleven EU countries.

Table 14.2. A selection of P4Q programmes in hospital care in Europe.

Table 14.2

A selection of P4Q programmes in hospital care in Europe.

14.3.1. Primary care

Table 14.1 provides an overview of the most important characteristics of 14 P4Q programmes in primary care in 13 European countries (Croatia, the Czech Republic, Estonia, France, Germany, Italy, Latvia, Lithuania, Poland, the Republic of Moldova, Portugal, Sweden and the United Kingdom (UK)). The first P4Q programme in primary care was introduced in 2001 in the context of disease management programmes in Germany, while the last was introduced in 2016 in Poland (see Table 14.1). Most P4Q programmes are implemented at the national level, but Germany, Italy and Sweden have regional P4Q programmes. About half of all programmes are mandatory, while the other half are voluntary.

All programmes have a strong focus on incentivizing quality in chronic and preventive care – with the exception of the one known programme in Italy, which only focuses on chronic care. All programmes include indicators that target improved effectiveness of care (for example, provision of certain services, compliance with guidelines, improved coordination and achievement of certain health outcomes). Only four programmes also include indicators that aim at improved responsiveness of care in terms of patient experience or patient satisfaction.

Quality indicators in most programmes focus on structures and processes of care but five countries also measure quality in terms of intermediate or final outcomes. Intermediate health outcomes, such as the achievement of a certain blood-pressure or a certain blood-glucose level in a pre-defined proportion of a patient population, have been the target of programmes in France, Latvia and the UK. In addition, a final outcome – i.e. reduced hospitalization in patients with chronic diseases – is included as an indicator in P4Q programmes in Latvia and Lithuania (Mitenbergs et al., 2012; Murauskiene et al., 2013). Furthermore, programmes in Portugal, Sweden and the UK reward outcomes of patient satisfaction or patient experience of care. The programme in Poland is the only known programme rewarding correct and timely diagnosis and timely treatment of cancer (OECD, 2016). Coordination efforts are rewarded in French, German, Italian and Swedish P4Q programmes, while practice organization and implementation of information technology and provision of other computer-based services are incentivized in at least seven countries, namely the Czech Republic, France, the Netherlands, Portugal, Spain, Sweden and the UK (Anell, Nylinder & Glenngård, 2012; OECD, 2016; Srivastava, Mueller & Hewlett, 2016). Finally, some programmes also reward improved access to care (for example, the scheme in the Czech Republic) – but this goes beyond the narrow definition of quality adopted by this book (see Chapter 1).

In all programmes, providers are rewarded with a bonus payment in relation to the measured quality of care – there are no penalties in any of the countries, except in certain regions of Sweden. The bonus is usually relatively small (<5% of total income) and is paid in relation to absolute performance. This means that the bonus of an individual provider is independent from the performance of other providers, except in certain regions of Sweden (Lindgren, 2014), where relative achievement compared to peers is rewarded. Only four programmes (in Croatia, France, Portugal and the UK) pay a bonus of more than 10%.

In Portugal bonuses are paid to physicians (up to 30% of income) and nurses (up to 10% of income) working in organizationally mature Family Health Units (FHU) that have gained greater autonomy from public administration (Biscaia & Heleno, 2017). Bonuses depend on achievements related to preventive and monitoring services in vulnerable populations (pregnant women, children, patients with diabetes or high blood-pressure) and in women of reproductive age (Almeida Simoes et al., 2017; Srivastava, Mueller & Hewlett, 2016).

Under the Quality and Outcomes Framework (QOF), implemented in the UK in 2004, practices could originally receive a bonus of up to 25% of income until 2013, when this share was reduced to 15% (Roland & Guthrie, 2016). The bonus comprises an up-front payment at the beginning of the year and achievement payments at the end. Points are awarded for the achievement of each incentivized indicator, and total payment depends on the monetary value of a QOF point, practice list size and prevalence data (NHS Digital, 2016). Indicators and the value of QOF points differ between England, Northern Ireland, Scotland and Wales. Initially, the scheme in England comprised 146 incentivized indicators from clinical, public health, organizational and patient experience domains (Doran et al., 2006; Gillam & Steel, 2013). However, in 2015 the number of indicators was reduced to 77; while many indicators were retired, some other indicators, such as smoking cessation and osteoporosis, were newly introduced (NHS Digital, 2016; NHS Employers, 2011). Even though the QOF has been implemented as a voluntary programme, participation rates have been very high, ranging from 96% to 99% (around 7 600 to 8 000) of eligible practices in England.

The French programme Rémunération sur objectifs de santé publique (ROSP) provides an incentive of up to 11% of the usual income to primary care physicians and in some cases to specialists (Cashin, 2014). The second programme in France, Expérimentations de no u veaux modes de remuneration (ENMR) applies a different incentive structure from other identified schemes in Europe. The scheme is comprised of basic and optional requirements, while the payment for each type of requirement consists of fixed and variable payment. In order to be able to participate in the programme, a provider has to fulfil basic requirements (Minister of Finance and Public Accounts/Minister of Social Affairs, Health and Women’s Rights, 2015). Overall, the payment depends on the achievements in three categories – access to healthcare, work in multiprofessional teams (which aims at better coordination of care), and implementation of computerized information systems. The scheme provides a bonus of up to 5% of the provider’s income, 60% of which can be paid in advance at the beginning of a period (Srivastava, Mueller & Hewlett, 2016).

14.3.2. Hospital care

Table 14.2 provides an overview of 13 P4Q programmes in nine European countries. The first P4Q programme in hospital care was introduced in Luxembourg in 1998 and the last of the included programmes was implemented in Norway in 2014 (see Table 14.2). Identified programmes are typically mandatory, implemented at the national level mainly in western European countries.

The focus of all programmes in hospitals is on acute care. The majority of programmes includes indicators that either target improved effectiveness of care (for example, performing surgery or initiating treatment within a pre-specified period of time) or patient safety (for example, avoidance of 30-day readmissions, wrong-side surgery and hospital-acquired conditions). Responsiveness in terms of patient experience or in terms of patient satisfaction is part of programmes in Denmark, Norway, Sweden, and the UK’s Advancing Quality (AQ) and Commissioning for Quality and Innovation (CQUIN).

Most P4Q programmes for hospitals have a stronger focus on outcomes and/or processes than P4Q programmes in primary care (where the focus is on structures). Only P4Q programmes for hospitals in Croatia, Denmark, France and Luxembourg include indicators for structures. Final health outcomes are only measured in Norway (for example, five-year survival rate for different cancer types, 30-day survival rates after hospital admission for hip fracture, AMI and stroke) and in Croatia (all-cause-mortality). Patient-reported health outcomes are measured in Advancing Quality (AQ) in the north-west of England (for example, quality of life), while patient safety outcomes are measured in the English “Non-payment for never-events” programme in terms of reduction of 14 never-events including wrong-side surgery, wrong implant/prosthesis, and retained foreign object post procedure (AQuA, 2017; NHS England Patient Safety Domain, 2015). Outcomes in terms of patient experience and patient satisfaction (for example, experience or satisfaction with waiting times) are rewarded by programmes in Denmark, Norway, Sweden and England (within AQ and CQUIN) (Anell, 2013; AQuA, 2017; Olsen & Brandborg, 2016).

Acute myocardial infarction (AMI), acute stroke, renal failure, hip fracture, and hip and knee replacement surgery are the main medical conditions targeted by programmes in France, Italy, Norway, Portugal, Sweden and the UK for process quality improvement. A few countries target additional conditions, such as cancer (Norway), diabetes (Sweden, UK), postpartum haemorrhage (France) and a few more in the UK. Indicators concern timely treatment (for example, surgical treatment of hip-fracture within 48 hours of admission, initiation of cancer treatment within 20 days), appropriate disease management (for example, medication at admission, discharge and during the stay, disease monitoring and diagnostic activities), and care coordination (for example, referrals to rehabilitation and primary care, plans for disease management, discharge summary sent within seven days).

Nine of the 13 identified programmes have penalties – either as a withhold of reimbursement (for example, non-payment schemes in the UK), as a payment adjustment of usual payment depending on performance (for example, CQUIN in the UK, programmes in Italy, Norway, Portugal and Sweden), or as a pre-defined fine if the targets are not met (for example, Journalauditindikatoren in Denmark) (Kristensen, Bech & Lauridsen, 2016). Some of the programmes have both penalties and bonuses (for example, schemes in Denmark, Portugal and CQUIN in the UK). In France, Luxembourg and the AQ scheme in the UK programmes rewarded providers with a bonus payment. The size of bonus payments or penalties is usually relatively small (<2% of total hospital income) and the payment is almost always made in relation to absolute performance. Only in France, Norway and Portugal does the payment depend on relative performance of providers compared to their peers. In most countries the bonus or penalty amounts to less than 2% of the total hospital budget. The scheme in Croatia is the only one where as much as 10% of a hospital’s revenue depends on a broader measure of performance including activity- and quality-based indicators (MSPY, 2016).

The earliest programme, the Incitant Qualité (IQ) in Luxembourg, was established with the aim to improve patient-centredness, and the sensibility of actors for quality of care. In the first four years the programme targeted prevention of nosocomial infections, implementation of electronic health records, preventive care and pain management, as well as the technical quality of mammography. The financial incentive currently amounts to up to 2% of the annual budget. The reward depends on the number of achieved points on a scale of 0 to 100 and the corresponding percentage with respect to all the available points (i.e. 0% for 0–10 points, 10% for 10–20 points and so on) (Sante.lu, 2015).

The Norwegian Quality-Based Financing (QBF) programme was introduced as a pilot among four regions in Norway and covers all public secondary care providers and also private hospitals with a contract with the Regional Health Authority (RHA) in Norway in January 2014. The rewards are paid to the four RHAs according to their performance and the performance of hospitals in the region measured by process, outcome and patient satisfaction indicators. While most indicators are measured on the hospital level, the five-year survival rates for cancer are measured on the regional level. The patient satisfaction results came from the National Patient Satisfaction Survey. The QBF rewards four types of performance: reporting quality, minimum performance level, best performance and best relative improvement in performance of RHA. The rewards are based on achieved points for the reporting quality and the three indicator types (outcome indicators – 50 000 points, process indicators – 20 000 points and indicators of patient satisfaction – 30 000 points). The fulfilment of reporting requirements is the prerequisite for the possibility to generate indicator-based points. QBF redistributes around 500 million Norwegian crones to RHAs according to the weighted performance of the regions and the regions’ hospitals. However, the RHAs have no fixed requirements regarding how to distribute the QBF rewards among regional hospitals (Olsen & Brandborg, 2016).

The French programme Incitation financière à l’amélioration de la qualité (IFAQ) was introduced as an experiment in 2012 and became a nationwide programme in 2016. The aim of the programme is to improve management of myocardial infarction, acute stroke, renal failure, the prevention and management of postpartum haemorrhage, documentation and efficient medication prescription. Only the upper 20% of the providers with the highest performance receive a bonus between 0.2 and 0.6% of total income. The total remuneration of the scheme amounts to between €15 000 and €500 000 (Minister of Social Affairs and Health, 2016).

14.4. The effectiveness and cost-effectiveness of pay for quality initiatives

The available evidence about the effectiveness of P4Q programmes has been summarized in 31 reviews published between 1999 and 2016. Tables 14.3, 14.4 and 14.5 provide an overview of the characteristics, methods and results of the included reviews. Most reviews were performed in the US and the UK. Seven reviews were conducted in non-English-speaking countries, and one review was in Portuguese. Nineteen reviews evaluated P4Q programmes in primary care (Table 14.3), nine1 reviews investigated effects in both primary and hospital care (Table 14.4) and only three reviews had an exclusive focus on hospital care (Table 14.5).

Table 14.3. Overview of systematic reviews evaluating P4Q schemes in primary care.

Table 14.3

Overview of systematic reviews evaluating P4Q schemes in primary care.

Table 14.4. Overview of systematic reviews evaluating P4Q schemes in both primary and hospital care.

Table 14.4

Overview of systematic reviews evaluating P4Q schemes in both primary and hospital care.

Table 14.5. Overview of systematic reviews evaluating P4Q schemes in hospital care.

Table 14.5

Overview of systematic reviews evaluating P4Q schemes in hospital care.

Five reviews focused solely on preventive care, while another three focused only on chronic care. Three reviews evaluated the effectiveness of P4Q in comparison to other interventions, with one review focusing on audit and feedback (Ivers et al., 2012), one review focusing on different interventions that can improve the appropriate use of imaging (French et al., 2010), and one review focusing on financial incentives (not only P4Q) for prescribers (Rashidian et al., 2015). In addition, two out of the three reviews reported by Damberg et al. (2014) evaluated accountable care organization models (ACOs) and bundled payment (BP) programmes, which aimed to improve quality and to simultaneously reduce costs of care.

The number of studies included in each review varies from two studies included by Giuffrida et al. (2000) to 128 studies included by van Herck et al. (2010). The original studies (around 400) included in the 31 reviews were conducted between the “early 1980s” (Armour et al., 2001) and 2013 (Milstein & Schreyoegg, 2016 suppl.), and reported between 1991 and 2015. Overall, reviews found the quality of included studies to be low to moderate. Most evidence stems from studies without a control group, i.e. studies of observational (for example, cross-sectional, longitudinal studies) and quasi-experimental nature (for example, uncontrolled before-after studies – UBA, time-series analyses). Even the relatively few available studies with a control group (approx. n = ≤100), such as randomized controlled trials (RCTs, n = ≤10), controlled before-after studies (CBA) and interrupted time-series (ITS) and other quasi-experimental designs with a control group, exhibit a number of biases.

With the exception of the reviews by Huang et al. (2013) and Ogundeji, Bland & Sheldon (2016), all the included systematic reviews synthesized included studies in a narrative manner. Ogundeji, Bland & Sheldon (2016) conducted a meta-analysis and a meta-regression, while Huang et al. (2013) only performed a meta-analysis.

14.4.1. Effectiveness of P4Q in primary care

The most frequently evaluated programme in primary care was QOF but most reviews evaluated a range of P4Q programmes in the US. Programmes in other European countries and in the Asia-Pacific region were evaluated only by individual studies included in the reviews (see Table 14.3).

The effectiveness of QOF has been evaluated by seven reviews in total, summarizing evidence from a total of 71 individual studies (Christianson, Leatherman & Sutherland, 2007, 2008; Gillam, Siriwardena & Steel, 2012; Hamilton et al., 2013; Houle et al., 2012; Kondo et al., 2015; Langdown & Peckham, 2014; Lin et al., 2015). The best evidence is available from five reviews that included studies evaluating at least four programme years after the start of the programme in 2004 and using results of ITS and other studies that accounted for secular trends (Gillam, Siriwardena & Steel, 2012; Houle et al., 2012; Kondo et al., 2015; Langdown & Peckham, 2014; Lin et al., 2015). Based on this body of evidence, review authors concluded that performance of primary care providers significantly improved in almost all process-of-care indicators (for example, smoking cessation activities, diabetes management activities) during the first year of the programme, with some improvements greater than 30 percentage points, while intermediate health outcomes (for example, blood pressure, cholesterol and blood glucose level under control) showed less improvement.

In subsequent years (2005 to 2007) performance reached a plateau but continued to slowly improve for process-of-care indicators in both chronic and preventive care (Gillam, Siriwardena & Steel, 2012; Houle et al., 2012; Kondo et al., 2015). However, this slow improvement was, in fact, very similar to the underlying trend before the implementation of QOF, and for some health outcomes (for example, blood pressure, cholesterol and blood glucose level under control), the observed improvement was even below the pre-QOF trend (Damberg et al., 2014; Houle et al., 2012; Kondo et al., 2015; Langdown & Peckham, 2014). In addition, no effect was observed on final health outcomes such as incidence of AMI, stroke, renal failure and all-cause mortality (Damberg et al., 2014; Kondo et al., 2015). In general, positive effects of QOF on process-of-care indicators were more often reported by observational studies and studies without a control group (Gillam, Siriwardena & Steel, 2012; Houle et al., 2012; Kondo et al., 2015), while effects on health outcomes were mixed and inconclusive.

Reported results of P4Q programmes in non-European countries are somewhat similar to those of QOF. Reviews identified 124 studies evaluating effects of P4Q programmes in primary care for chronic conditions. In general, short-term and observational or uncontrolled quasi-experimental studies frequently reported large positive effect sizes for process-of-care indicators in chronic care patients independent of the disease (Damberg et al., 2014; Houle et al., 2012; Kondo et al., 2015). Better designed studies, such as ITS, CBAs and other quasi-experimental designs with a comparison group, examining data over a longer time period (for example, several years before and several years after the implementation of a P4Q programme), found no effect or a slightly positive effect (Damberg et al., 2014; Houle et al., 2012; Kondo et al., 2015). Only small positive effects on chronic care management could be found in the networks within the ACOs (Damberg et al., 2014).

Reviews investigating effects of P4Q schemes on preventive care did not find convincing evidence for the effectiveness of P4Q interventions on preventive services (87 studies). Again, higher quality studies, i.e. those with an intervention and a control group, reported positive results only for individual process-of-care measures. For example, positive effects were found on colorectal and cervical cancer screening rates, on influenza immunization rates and on smoking cessation activities (for example, recording of smoking status and provision of cessation advice) (Damberg et al., 2014; Giuffrida et al., 2000; Hamilton et al., 2013; Houle et al., 2012; Kondo et al., 2015; Sabatino et al., 2008; Scott et al., 2011; Town et al., 2005; van Herck et al., 2010). However, no effects were found on screening rates for other cancer types, on screening referrals, as well as on adherence to cancer screening guidelines, and on paediatric immunization (Armour et al., 2001; Damberg et al., 2014; Sabatino et al., 2008; Town et al., 2005). Hamilton et al. (2013) identified seven studies which investigated effects of P4Q interventions on quit rates and smoking prevalence. One RCT and one cluster RCT found no superiority of interventions which applied a financial incentive for a healthcare provider over a control group or over other types of intervention on quit rates (Roski et al., 2003; Salize et al., 2009). In addition, the identified decrease of smoking prevalence could not be attributed to the P4Q intervention in the United Kingdom (QOF), nor in Taiwan (Hamilton et al., 2013).

Two reviews reported results of 11 studies that had investigated effects of P4Q on final health outcomes in non-European countries (Damberg et al., 2014; Kondo et al., 2015); seven of the 11 studies were of low quality and found positive effects on diabetes-related hospitalization and complications in the long-term, on reduced emergency department visits, on depression treatment response and on neonatal intensive care unit admissions. Two studies, one of good and one of low quality, found no effect on 30-day mortality, readmission, hospitalization and emergency department visits related to diabetes, AMI, heart failure and pneumonia (Damberg et al., 2014). For the remaining two studies, reviews reported detrimental effects on acute emergency department visits related to asthma, diabetes and heart failure (Kondo et al., 2015).

Effects of P4Q on responsiveness of care are reported in five reviews. Gillam, Siriwardena & Steel (2012) found on the basis of six observational studies of patient experience in QOF that no statistically significant changes in communication, nursing care, coordination or overall satisfaction were reported by patients between 2003 and 2007. However, the same six original studies found that timely access to chronic care worsened in terms of continuity of care and visits to the usual physician, but not in terms of urgent appointments, which actually improved statistically significantly. In general, and especially for older patients, access to care in QOF worsened. Christianson, Leatherman & Sutherland (2007) and van Herck et al. (2010) reported for several international P4Q programmes that patient satisfaction with care did not change. Two other reviews highlighted that positive effects on patient experience reported by original studies could not be clearly attributed to a P4Q programme, either because of structural changes implemented as part of the programme (for example, implementation of electronic reminder and prescribing systems) or because other quality improvement interventions were implemented simultaneously with the P4Q programme (Damberg et al., 2014; Kondo et al., 2015).

Finally, one cluster-RCT identified by Ivers et al. (2012) evaluated the effects of financial incentives compared to audit and feedback on test-ordering. The financial incentives turned out to be less effective than audit and feedback in reducing test ordering.

14.4.2. Effectiveness of P4Q in hospital care

Reviews that evaluated programmes in hospital care (Tables 14.4 and 14.5) identified 30 studies of 15 P4Q programmes, most of which were located in the US and incentivized primarily process-of-care measures (Armour et al., 2001; Barreto, 2015; Christianson, Leatherman & Sutherland, 2007, 2008; Damberg et al., 2007, 2014; Kondo et al., 2015; Korenstein et al., 2016; Mehrotra et al., 2009; Milstein & Schreyoegg, 2016). P4Q programme effects on health outcomes were evaluated by 13 studies. Only few programmes were evaluated exhaustively – such programmes are “Advancing Quality” in the UK evaluated by four studies and the discontinued HQID (2003–2009) in the US evaluated by 17 studies.

Reviews reported that studies with a comparison group found predominantly small short-term and often statistically non-significant positive effects on a composite score that combined several process-of-care measures, or positive effects on individual process-of-care indicators (Damberg et al., 2007, 2014; Kondo et al., 2015; Mehrotra et al., 2009; Milstein & Schreyoegg, 2016). Highly positive effects were identified in the initial phase of HQID, while in the long term the effects were not sustained (Damberg et al., 2014; Mehrotra et al., 2009; Milstein & Schreyoegg, 2016). In contrast, the positive effects of the initial phase of the more recent Hospital Value-Based purchasing incentive Payment programme (HVBP) were not statistically significant (Kondo et al., 2015; Milstein & Schreyoegg, 2016). In three US programmes (MassHealth, Non-payment for HACs and Baylor Healthcare System) evaluated by three studies with relatively strong designs (i.e. with a comparison group or with time-trend adjustment), positive programme effects were observed only on individual process-of-care measures related to pneumonia, AMI and CHF management (for example, influenza vaccination in pneumonia patients – one out of the 19 pneumonia measures) (Damberg et al., 2014; Kondo et al., 2015). Six studies with no comparison group found positive effects on breast cancer, AMI and CHF management, on obstetric services and common surgeries (Armour et al., 2001; Damberg et al., 2007, 2014; Kondo et al., 2015; Mehrotra et al., 2009). One UBA evaluation of a P4Q programme in Taiwan found no effect on tuberculosis treatment length (Kondo et al., 2015).

Similar results were also found with respect to health outcomes. The rate of decrease of risk-adjusted mortality associated with AMI, heart failure or pneumonia was larger in the initial phase of Advancing Quality than in the long term. That is, 42 months after the introduction of the programme, no further improvements in mortality rates were observed and hospitals in other regions of England showed greater reductions in mortality (Damberg et al., 2014; Kondo et al., 2015; Milstein & Schreyoegg, 2016). The effects of other P4Q programmes were mixed. Positive effects were identified for different types of health outcomes: five-year breast cancer survival, negative surgical margins and breast cancer recurrence rate in the Taiwanese Breast Cancer Pay for Performance programme (BC-P4P), nine-months tuberculosis cure rate in the Taiwanese Tuberculosis Pay for Performance (TB-P4P) programme and on quality-adjusted life-years (QALYs) associated with AMI and CHF in the Blue Cross Blue Shield Michigan P4P (BCBS-P4P) programme (Christianson, Leatherman & Sutherland, 2007, 2008; Damberg et al., 2007, 2014; Kondo et al., 2015; Mehrotra et al., 2009; Milstein & Schreyoegg, 2016; van Herck et al., 2010). However, in the original Taiwanese studies, no information on study design was provided, while other studies either lacked a comparison group (for example, BCBS-P4P), or lacked adjustment for time-trend and the coincident public-reporting effects (for example, the Italian DRG-P4P) (Kondo et al., 2015; Mehrotra et al., 2009). In three studies Damberg et al. (2014) and Mehrotra et al. (2009) found no difference between HQID hospitals and the comparison group in mortality rates associated with AMI, CHF and pneumonia.

Patient safety or utilization outcomes were evaluated by seven studies included in six reviews with respect to readmissions, length-of-stay (LOS), surgery-related complications or infections, blood catheter-associated infections and other hospital acquired conditions (HACs) in seven programmes –Advancing Quality; Hawaii Medical Service Association Hospital Pay for Performance (HMSA-P4P); HQID; HVBP; Non-payment for HACs by the US Centers for Medicare and Medicaid Services; Geisinger ProvenCareSM integrated delivery system; and MassHealth P4Q. Positive and statistically significant effects on preventable conditions or LOS were only identified by two studies in HMSA-P4P and in Non-payment for HACs, while in four studies positive effects were small and statistically not significant (Christianson, Leatherman & Sutherland, 2008; Damberg et al., 2014; Korenstein et al., 2016; Mehrotra et al., 2009; Milstein & Schreyoegg, 2016).

Responsiveness in terms of patient experience was evaluated by four studies in five reviews. The reviews by Kondo et al. (2015) and Milstein & Schreyoegg (2016) did not find evidence for improved patient experience of care after the introduction of HVBP but rather found a statistically non-significant worsening of care. Patient satisfaction with inpatient care in HMSA-P4P hospitals improved by a few percentage points. However, the evaluation did not involve a control group and the statistical significance was not calculated either (Christianson, Leatherman & Sutherland, 2008; Damberg et al., 2014; Mehrotra et al., 2009).

14.4.3. Cost-effectiveness

Emmert et al. (2012) is the only review that examined economic evaluations of P4Q programmes. It identified only three full economic evaluations. Six studies were partial economic evaluations, which evaluated costs and consequences separately or assessed only the impact on costs. The reviews by Christianson, Leatherman & Sutherland (2007), van Herck et al. (2010), Gillam, Siriwardena & Steel (2012), Hamilton et al. (2013) and Kondo et al. (2015) identified three other studies with partial economic evaluations.

All full economic evaluations included in the review by Emmert et al. (2012) reported positive cost-effectiveness. All three studies evaluated the effects of financial incentives on processes of care in primary or hospital care in the US. The RCTs by Kouides et al. (1998) evaluated effects of additional bonuses on influenza immunization coverage. The study found additional costs of $4 362 and $1 443 for additional immunizations. Overall, in the intervention group median improvement of coverage was 10.3% compared to the pre-intervention period, while in the control group median improvement was only 3.5%. The RCT by An et al. (2008) evaluated effects of incentives on referrals and enrolment in a quit smoking programme. The programme resulted in 1 483 total referrals and $95 733 total costs ($64 per referral) in the intervention group and 441 total referrals and $8937 total costs in the control ($20 per referral) group. The referrals in the intervention group resulted in 289 additional enrolees in the quit smoking programme and $300 per additional enrolee. The study by Nahra et al. (2006) evaluated the hospital BCBS-P4P programme, focusing on effects for AMI and CHF patients, and estimated costs per QALYs gained of between $12 967 and $30 081.

Most partial economic evaluations also reported positive results (Emmert et al., 2012; van Herck et al., 2010). Only one cost-effectiveness study conducted by Salize et al. (2009) evaluated the effects side using a health outcome, i.e. smoking abstinence. The RCT compared three arms with different combinations of interventions – physician training, financial incentive and free medication prescription – to usual care. In contrast to the two arms containing free medication prescription, the combination of physician training and financial incentive turned out to be not cost-effective when comparing the intervention costs per smoking-abstinent patient to the usual treatment. Even the third arm, which contained training, free medication prescription and financial incentive, did not dominate over the arm containing only training and free medication prescription (Hamilton et al., 2013; Scott et al., 2011; van Herck et al., 2010).

In general, the economic evaluations included in identified reviews have a number of weaknesses: included analyses predominantly considered process-of-care indicators on the effects side and costs from the third-party-payer’s perspective on the costs side. Costs from the provider’s perspective, such as administrative costs or costs for participating in other quality improvement initiatives, were not taken into account, and the costs were rarely described in detail (Emmert et al., 2012). In addition, designs of the included analyses have several limitations (for example, lack of separation of the effects generated by public reporting, small sample sizes, unit-of-analysis errors, etc.), which restrict the reliability of their conclusions on cost-effectiveness (Emmert et al., 2012; Mehrotra et al., 2009). Furthermore, a number of evaluated programmes (for example, HQID, QOF and HVBP) has been found to be ineffective in the long term (Gillam, Siriwardena & Steel, 2012; Houle et al., 2012; Kondo et al., 2015). Therefore, cost-effectiveness, if any, could have only been achieved in the programme’s short term, when the combination of health gain and the sum of additional costs (administrative and reward costs) of the programme did not exceed a pre-specified amount. For many P4Q programmes, reviews found no positive effects which means that these programmes could not be cost-effective because they required additional financial resources.

14.5. How can pay for quality programmes be implemented? What are the organizational and institutional requirements?

The implementation of P4Q schemes is quite complex as many strategic and technical questions need to be addressed. Eijkenaar (2013) has proposed three broad strategic questions that need to be considered and we have added another two:

  1. What to incentivize?
  2. How to measure quality?
  3. Whom to incentivize?
  4. How to incentivize? and
  5. How to implement and administer a P4Q programme?

14.5.1. What to incentivize?

The question of “what to incentivize?” requires scrutiny of the quality objectives that the payer wishes to prioritize. It is important that programmes focus on areas of quality where change is needed, rather than on areas where performance is already widely embedded in clinical practice (Lin et al., 2015; van Herck et al., 2010). Piecemeal attention to only some aspects of quality might encourage neglect of non-incentivized aspects. Therefore, P4Q should likely embrace a comprehensive approach and aim at covering most (or many) relevant areas of care (for example, Milstein & Schreyoegg, 2016).

Furthermore, a common theme in the literature (for example, Doran, Maurer & Ryan, 2017; Roland & Dudley, 2015; van Herck et al., 2010) is that incentivized activities should be aligned with widely held professional principles. This is one of the reasons why P4Q schemes should be developed in collaboration with healthcare professionals. Schemes are unlikely to be effective unless they reinforce professional norms and beliefs. In fact, the principles of P4Q can be considered to be somewhat antithetic to the principles of professional practice, which imply doing the best for patients irrespective of financial reward. Therefore, the very existence of a P4Q scheme may be a signal that some aspects of current professional practice are unacceptable.

14.5.2. How to measure quality?

Indicators and metrics to be used as the basis of reward should be reliable and timely, and not vulnerable to distortion (such as the provision of high-quality care only to healthier patients) or mis-reporting (such as only reporting values desired by the scheme). Indicators may reflect the structures, processes or outcomes of quality, and the choice of indicators will involve a trade-off between on the one hand the simplicity and practicality of structural and process metrics, and on the other hand the greater relevance but also greater complexity of outcome measures (see also Chapter 3).

While the structures of care reflect provider characteristics, such as the qualifications or accreditation of staff, the link of such structures to the eventual desired quality outcomes is often quite remote. Structural indicators can be considered in P4Q programmes when other data collection is infeasible, or there is a clear link from the structure to eventual quality.

Rewarding the processes of care is a more direct approach towards promoting quality, so long as the incentivized metrics are known to be associated with the desired quality outcomes. Process-based schemes are the most practical approach towards P4Q in many circumstances, as they obviate the need to directly measure outcomes. They can be aligned with clinical guidelines to motivate professionals to adopt best practices (see also Chapter 9), especially if this requires changes to existing methods and investments such as retraining. In principle, it should be unnecessary to reward processes that are already embedded in good professional practice.

Rewarding the outcomes of care seeks to directly reward the desired results of high-quality care. Examples might include future health status, or health service utilization metrics, such as hospital readmission. However, although directly addressing health system objectives, outcome-related P4Q is also the most challenging type of scheme. Levels of quality attained may be highly dependent on the characteristics of the patients treated, so some sort of casemix adjustment to the performance metrics is essential in order to avoid cherry-picking of healthier patients (see Chapter 3). Methods of risk-adjustment can vary from crude approaches (for example, excluding “complex” patients from the calculations) to statistically sophisticated methods. Many authors argue that statistical risk-adjustment is preferable because excluding complex patients from the calculation means that quality of care for these patients will not be incentivized by the P4Q scheme (Christianson, Leatherman & Sutherland, 2007, 2008; Gillam, Siriwardena & Steel, 2012; Houle et al., 2012; Kondo et al., 2015; Langdown & Peckham, 2014). Another approach to deal with the potential risk for risk-selection can be to pay more for target achievement amongst patients with comorbidities, for example, achievement of blood pressure control in diabetic patients or among patients with chronic kidney disease, because these targets are more difficult to achieve (Roland & Dudley, 2015). Such stronger incentives would be desirable also from a societal perspective since they may prevent costly disease-related complications in the long term.

Another challenge for outcome-based metrics is that some aspects of high-quality care may take a long time to materialize, rendering them infeasible as a basis for measurement and reward. Therefore, although they offer the most direct link to desired objectives, the outcome-focused approach towards P4Q is likely to have limited applicability in practice.

14.5.3. Who to incentivize?

The question of “who to incentivize?” is often a finely balanced decision (Conrad, 2015; Kondo et al., 2015, 2016; Rynes, Gerhart & Parks, 2005). It is generally easiest administratively for the payer to target entire provider organizations. However, this requires that the organizations have some leverage over the practitioners on whom most aspects of quality ultimately depend. In contrast, targeting clinical teams or individual practitioners may sacrifice the collective responsibility and peer pressure needed to improve some aspects of quality, and the associated metrics may be less reliable and vulnerable to random fluctuation.

It is usually preferable to make participation in a P4Q scheme compulsory, especially as providers with unsatisfactory performance are the target of many schemes. Voluntary participation may result in a joining-in of already high-performing providers, which leads to a reward of the historical and not the improvement of performance (Christianson, Leatherman & Sutherland, 2007; Mehrotra et al., 2009). Secondly, depending on the scheme design, voluntary participation may prevent poorly performing providers from joining and may allow premature cancellation of participation in the programme (Scott et al., 2011). However, there may be circumstances when voluntary participation is needed to secure acceptance of the principle of P4Q, and if necessary the rewards can be designed to encourage high levels of participation.

14.5.4. How to incentivize?

A great variety of approaches towards incentivizing mechanisms have been tested and discussed in the literature (for example, Conrad, 2015; Doran, Maurer & Ryan, 2017; Kondo et al., 2015; Milstein & Schreyoegg, 2016; Roland & Dudley, 2015). Decisions must be taken on a wide range of characteristics of the quality-related payments, which are listed in Box 14.3. Choices will depend on criteria such as the disease area, the information available, administrative feasibility, the funds available, and the capacity of the payer and providers. Each of the decisions taken may influence the effects of the programme.

Box Icon

Box 14.3

Aspects of financial incentives that must be considered when planning a P4Q programme.

When deciding about the size of the financial incentive, prospective expected incremental costs of the quality improvement and the share of total provider’s income affected should be taken into account. If incentives are too small, they are likely to be ineffective, while very large incentives are unlikely to be cost-effective. For instance, Ogundeji, Bland & Sheldon (2016) showed in a meta-regression that the positive effects of a programme tend to be higher in programmes applying larger incentives (≥ 5% of annual income).

The decisions on the structure of the financial incentive (for example, reward vs. penalty), the source of the payment (for example, “old” money – withholding part of the annual payment at the start of a period and redistributing it according to performance at the end of the period, or “new” money – payment of additional bonuses), the payment basis (for example, absolute vs. relative measurement, attainment vs. improvement) and performance targets (for example, single elements vs. composite score, availability of a threshold) influence the reaction of providers to the financial incentive. Each of these elements considered individually has various advantages and disadvantages.

Rewards of absolute performance measures are easy to manage and they provide some certainty of payment to providers. However, evidence from many programmes shows that absolute performance rewards often do not lead to the desired effects in the long term. The predetermined absolute performance thresholds hamper continuous incentives for further improvement of quality in healthcare if the targets are not revised on a regular basis (Langdown & Peckham, 2014).

There are also numerous negative aspects of penalties and relative performance measurements (Arnold, 2017; Conrad, 2015). They may lead to discrimination and unfairness and result in low acceptance and negative (unintended) behavioural reactions of providers or professionals. However, relative measures can incentivize continuous improvement and penalties usually have a stronger influence on performance due to the loss aversion of individuals (Emanuel et al., 2016). Individuals will make more effort to protect their revenues rather than to earn an uncertain reward. Furthermore, redistribution of “old” money can be perceived as unfair by providers (Milstein & Schreyoegg, 2016), which may again result in negative reactions (Eijkenaar, 2013; Kahneman, Knetsch & Thaler, 1986).

There is no clear evidence that would support the superiority of one incentive structure over another. However, blended payment systems, combining various characteristics, can reduce the unintended consequences. For example, the combination of “old” and “new” money, as well as of rewards, penalties and relative performance measures, can exploit the advantages of these elements, while avoiding some of the disadvantages. Loss aversion of individuals can be exploited by rewarding P4Q participants with part of a quality-related payment at the beginning of a period, which will be adjusted for performance at the end of the period. Another approach can be to fine providers who are not achieving quality aims, while a bonus is paid if further performance goals are reached.

In general, the emphasis of P4Q programmes should be to reward improvement of individual performance from previous levels, especially compared to the previous period. Highly competitive approaches that reward only the top 20% of providers with the highest performance or the largest improvement should rather be avoided because of the aforementioned potential negative consequences. However, whichever choices are made, it is important that they are codified very clearly, with statements of entitlements, conditions, time horizons and criteria for receipt of funds.

14.5.5. How to implement and administer?

In order to increase acceptance of a P4Q programme, all relevant stakeholders (providers, patients and payers) should be involved from the beginning of programme development, through implementation and evaluation (Damberg et al., 2014; van Herck et al., 2010). When implementing a programme, participating providers have to be trained about involved measures and about the relationship between the measures and the financial incentives (Kane et al., 2004; Kondo et al., 2015; Milstein & Schreyoegg, 2016; Sorbero et al., 2006). Time horizons of financial incentives should be clearly communicated, and allocation of rewards within an institution should be clear, too. Commissioners of a programme should assume that all participating providers can achieve the pre-specified targets in a short period of time and calculate funds accordingly. Furthermore, it is important that all relevant aspects of quality are monitored – not only incentivized aspects – even if they are not included in the P4Q scheme.

Finally, implemented programmes have to be monitored and evaluated on a regular basis. A number of recommendations for P4Q evaluations emerge from the available literature (for example, Damberg et al., 2014; Kondo et al., 2015; Mehrotra et al., 2009; Milstein & Schreyoegg, 2016). Evaluations should usually be planned before a P4Q programme starts and an appropriate evaluation design selected, depending on the number of participating providers and the time horizon of the programme. For programmes with high participation rates (for example, almost all hospitals), it is appropriate to apply an interrupted time-series design when assessing programme effectiveness. In doing so, performance and quality data should be collected for several years before and after the implementation of the programme. However, because studies without a comparison group systematically over-estimate the positive effects of P4Q programmes (Ogundeji, Bland & Sheldon, 2016), evaluation designs should, ideally, contain a comparison group, adjust for baseline performance of participating and non-participating providers, and account for secular trends. Furthermore, an evaluation should account for the implementation of concurrent quality improvement interventions, such as audit and feedback and public reporting, and also for the – often – frequent changes in programme design.

14.6. Conclusions for policy-makers

For obvious reasons, P4Q is not a panacea for solving a health system’s quality problems. Despite the many implemented programmes in Europe, and even more programmes in the United States, the effectiveness and cost-effectiveness of P4Q programmes remain unclear. However, implementing P4Q programmes is complex and the main lessons concerning the design of P4Q programmes are summarized in Box 14.4.

Box Icon

Box 14.4

Conclusions with respect to P4Q programme design.

Our review of existing P4Q schemes in Europe found 27 programmes in 16 European countries, with 14 programmes in primary care and 13 programmes in hospital care. Most P4Q programmes in primary care focus on quality in terms of structures and processes. Programmes for hospitals also focus on quality of processes but they focus just as often on quality of outcomes. Regardless of the increasing number of programmes in Europe, available evidence about the effectiveness of P4Q mostly stems from the United States or from England. P4Q programmes in other European countries have rarely been evaluated.

Reviews of P4Q programmes in primary care showed that incentivizing process quality more often had a positive effect than incentivizing intermediate health outcomes. In contrast, P4Q programmes in hospital care appeared to be ineffective with respect to process quality, while the evidence on their effectiveness with regard to final health outcomes and patient safety indicators was inconclusive. Effects on final health outcomes were rarely evaluated in primary care and available results were partly contradictory. The relationship between intermediate and final health outcomes also remains unclear for evaluated P4Q programmes. Patient satisfaction and patient experience in primary care did not improve and sometimes deteriorated with respect to continuity of care, communication, nursing care, coordination and overall care satisfaction. The effect of hospital P4Q programmes on patient experience and patient satisfaction was rarely evaluated and showed only minor changes. However, patient satisfaction and patient experience are important indicators that should not be disregarded.

A few evaluations showed that P4Q programmes were less effective compared to other quality improvement interventions, such as public reporting, and audit and feedback. Cost-effectiveness of P4Q was rarely evaluated. Two studies were found that show P4Q interventions to be cost-effective from a third-party-payer perspective. However, these results need to be viewed in the context of the larger body of literature that found no or minor effects on improved quality of care. As programmes certainly entail additional costs, it is rather unlikely that these programmes are cost-effective.

Even if there is limited evidence about the effectiveness of P4Q programmes, there is substantial evidence from various countries that implementing such programmes is complex. A number of important governance issues must be resolved for any P4Q scheme to function properly. The most basic is that arrangements must be put in place to develop the content and structures of the scheme, and to review and update the quality metrics. Involvement of relevant professionals and patients is important, but the interests of payers must also be protected.

A fundamental element of any P4Q scheme is the information on which its payments are based, including any information used for risk-adjustment. Furthermore, proper monitoring requires information on certain non-incentivized aspects of care, to ensure that they have not been harmed by the P4Q scheme. It is likely that receipt of funds should be conditional on timely provision of relevant data by the providers involved, and that the quality of the data should be properly monitored and validated. More generally, the payer should have the capacity to monitor adverse behavioural responses on the part of providers, such as “cream-skimming” healthier patients.

Finally, any P4Q scheme should be subjected to routine monitoring and evaluation. This should seek to identify the benefits of the scheme and any adverse consequences. Payers may consider some sort of phased introduction, so that the scheme can be properly evaluated. The contents of the scheme should be regularly reviewed and refreshed, as certain elements are likely to become redundant (for example if variations in performance are reduced) and new concerns arise.

References

  • Achat H, McIntyre P, Burgess M. Health care incentives in immunisation. Australian and New Zealand Journal of Public Health. 1999;23(3):285. [PubMed: 10388173]
  • Almeida Simoes J de, et al. Portugal: Health System Review. Health Systems in Transition. 2017;19(2) [PubMed: 28485714]
  • An LC, et al. A randomized trial of a pay-for-performance program targeting clinician referral to a state tobacco quitline. Archives of In ternal Medicine. 2008;168(18):1993. [PubMed: 18852400]
  • Anell A. Vårdval i specialistvården: Utveckling och utmaningar. Stockholm: Sveriges kommuner och landsting; 2013.
  • Anell A, Nylinder P, Glenngård AH. Vårdval i primärvården: Jämförelse av uppdrag, ersättningsprinciper och kostnadsansvar. Stockholm: Sveriges kommuner och landsting; 2012.
  • AQuA. About Us: How does Advancing Quality measure performance? 2017. Available at: http://www​.advancingqualitynw​.nhs.uk/about-us/, accessed 26 October 2017.
  • Armour BS, et al. The effect of explicit financial incentives on physician behavior. Archives of Internal Medicine. 2001;161(10):1261. [PubMed: 11371253]
  • Arnold DR. Countervailing incentives in value-based payment. Healthcare (Amsterdam, Netherlands). 2017;5(3):125. [PubMed: 28822499]
  • Barreto JO. [Pay-for-performance in health care services: a review of the best evidence available] Cien Saude Colet. 2015;20(5):1497. [PubMed: 26017951]
  • Biscaia AR, Heleno LCV. A Reforma dos Cuidados de Saúde Primários em Portugal: portuguesa, moderna e inovadora. Ciencia & saude coletiva. 2017;22(3):701. [PubMed: 28300980]
  • Busse R, Blümel M. Payment systems to improve quality, efficiency and care coordination for chronically ill patients – a framework and country examples. In: Mas N, Wisbaum W (eds.), editors. The “Triple Aim” for the future of health care. Madrid: Spanish Savings Banks Foundation (FUNCAS); 2015.
  • Cashin C. European Observatory on Health Systems and Policies series. Maidenhead, England: Open University Press, McGraw-Hill Education; 2014. Paying for performance in health care: Implications for health system performance and accountability.
  • Christianson JB, Knutson DJ, Mazze RS. Physician pay-for-performance. Implementation and research issues. Journal of General Internal Medicine. 2006;21 Suppl 2:S9–S13. [PMC free article: PMC2557129] [PubMed: 16637965]
  • Christianson JB, Leatherman S, Sutherland K. Financial incentives, healthcare providers and quality improvements: a review of the evidence. London: Health Foundation; 2007.
  • Christianson JB, Leatherman S, Sutherland K. Lessons from evaluations of purchaser pay-for-performance programs: a review of the evidence. Medical Care Research and Review. 2008;65(6) Suppl:5S–35S. [PubMed: 19015377]
  • Conrad DA. The Theory of Value-Based Payment Incentives and Their Application to Health Care. Health Services Research. 2015;50 Suppl 2:2057. [PMC free article: PMC5338202] [PubMed: 26549041]
  • Conrad DA, Christianson JB. Penetrating the “black box”: financial incentives for enhancing the quality of physician services. Medical Care Research and Review. 2004;61(3) Suppl:37S–68S. [PubMed: 15375283]
  • Damberg CL, et al. An Environmental Scan of Pay for Performance in the Hospital Setting: Final Report. Washington, DC: RAND Corporation; 2007.
  • Damberg CL, et al. Measuring Success in Health Care Value-Based Purchasing Programs: Findings from an Environmental Scan, Literature Review, and Expert Panel Discussions. Washington, DC: RAND Corporation; 2014. [PMC free article: PMC5161317] [PubMed: 28083347]
  • De Bruin SR, Baan CA, Struijs JN. Pay-for-performance in disease management: a systematic review of the literature. BMC Health Services Research. 2011;11:272. [PMC free article: PMC3218039] [PubMed: 21999234]
  • Doran T, Maurer KA, Ryan AM. Impact of Provider Incentives on Quality and Value of Health Care. Annual Review of Public Health. 2017;38:449. [PubMed: 27992731]
  • Doran T, et al. Pay-for-performance programs in family practices in the United Kingdom. New England J ournal of Medicine. 2006;355(4):375. [PubMed: 16870916]
  • Dudley RA, et al. The Impact of Financial Incentives on Quality of Health Care. Milbank Quarterly. 1998;76(4):649. [PMC free article: PMC2751095] [PubMed: 9879306]
  • Dudley RA, et al. Strategies To Support Quality-based Purchasing: A Review of the Evidence. AHRQ Publication, No. 04-0057. Rockville MD: Agency for Healthcare Research and Quality; 2004. [PubMed: 20734506]
  • Eckhardt H, et al. Effectiveness and cost-effectiveness of pay for quality initiatives in high-income countries: a systematic review of reviews. PROSPERO 2016: CRD42016043043. 2016. Available at: http://www​.crd.york.ac​.uk/PROSPERO/display_record​.asp?ID=CRD42016043043, accessed 17 May 2019.
  • Eijkenaar F. Key issues in the design of pay for performance programs. European Journal of Health Economics. 2013;14(1):117. [PMC free article: PMC3535413] [PubMed: 21882009]
  • Emanuel EJ, et al. Using Behavioral Economics to Design Physician Incentives That Deliver High-Value Care. Annals of Internal Medicine. 2016;164(2):114. [PubMed: 26595370]
  • Emmert M, et al. Economic evaluation of pay-for-performance in health care: a systematic review. European Journal of Health Economics. 2012;13(6):755. [PubMed: 21660562]
  • FHL. Le modèle des Incitants Qualité – Bilan des démarches communes EHL – CNS et perspectives. Fédération des Hôpitaux Luxembourgeois. 2012
  • French SD, et al. Cochrane Database of Systematic Reviews. 1. West Sussex: John Wiley & Sons, Ltd; 2010. Interventions for improving the appropriate use of imaging in people with musculoskeletal conditions; p. CD006094. [PMC free article: PMC7390432] [PubMed: 20091583]
  • Gillam S, Siriwardena AN, Steel N. Pay-for-performance in the United Kingdom: impact of the quality and outcomes framework: a systematic review. Annals of Family Medicine. 2012;10(5):461. [PMC free article: PMC3438214] [PubMed: 22966110]
  • Gillam S, Steel N. The Quality and Outcomes Framework – where next. BMJ. 2013;346(2):f659. [PubMed: 23393112]
  • Giuffrida A, et al. Target payments in primary care: effects on professional practice and health care outcomes. Cochrane Database of Systematic Reviews. 2000;(3):CD000531. [PMC free article: PMC7032683] [PubMed: 10908475]
  • Hamilton FL, et al. Effectiveness of providing financial incentives to healthcare professionals for smoking cessation activities: systematic review. Tobacco Control. 2013;22(1):3. [PubMed: 22123941]
  • Houle SK, et al. Does performance-based remuneration for individual health care practitioners affect patient care? A systematic review. Annals of Internal Medicine. 2012;157(12):889. [PubMed: 23247940]
  • Huang J, et al. Impact of pay-for-performance on management of diabetes: a systematic review. Journal of Evidence-Based Medicine. 2013;6(3):173. [PubMed: 24325374]
  • Ivers N, et al. Audit and feedback: effects on professional practice and healthcare outcomes. Cochrane Database of Systematic Reviews. 2012;(6):CD000259. [PubMed: 22696318]
  • Kahneman D, Knetsch JL, Thaler R. Fairness as a Constraint on Profit Seeking: Entitlements in the Market. American Economic Review. 1986;76(4):728.
  • Kane RL, et al. Economic incentives for preventive care. Evidence Report/Technology Assessment (Summary). 2004;(101):1. [PMC free article: PMC4781426] [PubMed: 15526397]
  • Kondo K, et al. Understanding the Intervention and Implementation Factors Associated with Benefits and Harms of Pay for Performance Programs in Healthcare. Washington, DC: Department of Veterans Affairs; 2015. [PubMed: 27054229]
  • Kondo KK, et al. Implementation Processes and Pay for Performance in Healthcare: A Systematic Review. Journal of General Internal Medicine. 2016;31 Suppl 1:61. [PMC free article: PMC4803682] [PubMed: 26951276]
  • Korenstein D, et al. Do Health Care Delivery System Reforms Improve Value? The Jury Is Still Out. Medical Care. 2016;54(1):55. [PMC free article: PMC4869989] [PubMed: 26492216]
  • Kouides RW, et al. Performance-based physician reimbursement and influenza immunization rates in the elderly. The Primary-Care Physicians of Monroe County. American Journal of Preventive Medicine. 1998;14(2):89. [PubMed: 9631159]
  • Kristensen SR, Bech M, Lauridsen JT. Who to pay for performance? The choice of organisational level for hospital performance incentives. European Journal of Health Economics. 2016;17(4):435–42. [PubMed: 25860814]
  • Kronick R, Casalino LP, Bindman AB. Introduction. Apple Pickers or Federal Judges: Strong versus Weak Incentives in Physician Payment. Health Services Research. 2015;50 Suppl 2:2049. [PMC free article: PMC5338199] [PubMed: 26769059]
  • Langdown C, Peckham S. The use of financial incentives to help improve health outcomes: is the quality and outcomes framework fit for purpose? A systematic review. Journal of Public Health (Oxford, England). 2014;36(2):251. [PubMed: 23929885]
  • Lin Y, et al. Impact of Pay for performance on Behavior of Primary Care Physicians and Patient Outcomes. Journal of Evidence-Based Medicine. 2015;9(1):8–23. [PubMed: 26667492]
  • Lindgren P. Ersättning i sjukvården: Modeller, effekter, rekommendationer. Stockholm: SNS Förl; 2014.
  • Mehrotra A, et al. Pay for performance in the hospital setting: what is the state of the evidence. American Journal of Medical Quality. 2009;24(1):19. [PubMed: 19073941]
  • Milstein R, Schreyoegg J. Pay for performance in the inpatient sector: a review of 34 P4P programs in 14 OECD countries. Health policy (Amsterdam, Netherlands). 2016;120(10):1125–40. [PubMed: 27745916]
  • Minister of Finance and Public Accounts/Minister of Social Affairs, Health and Women’s Rights. Arrêté du 23 février 2015 portant approbation du règlement arbitral applicable aux structures de santé pluri-professionnelles de proximité 2015:49.
  • Minister of Social Affairs and Health. Arrêté du 5 août 2016 fixant les modalités de calcul du montant de la dotation allouée aux établissements de santé en application de l’article L. 2016:162-22-20.
  • Mitenbergs U, et al. Latvia: Health system review. Health Systems in Transition. 2012;14(8) [PubMed: 23579000]
  • MSPY. National social report of Republic of Croatia. Ministry of Social Policy and Youth. 2016.
  • Murauskiene L, et al. Lithuania: Health System Review. Health Systems in Transition. 2013;15(2) [PubMed: 23902994]
  • Nahra TA, et al. Cost-effectiveness of hospital pay-for-performance incentives. Medical Care Research and Review. 2006;63(1) Suppl:49S–72S. [PubMed: 16688924]
  • NHS Digital. Quality and Outcomes Framework – Prevalence, Achievements and Exceptions Report: England, 2015–2016. 2016. Available at: http://www​.content.digital​.nhs.uk/catalogue​/PUB22266/qof-1516-rep-v2.pdf, accessed 13 March 2017.
  • NHS Employers. Quality and Outcomes Framework for 2012/13: Guidance for PCOs and practices. 2011. Available at: http://www​.nhsemployers​.org/your-workforce​/primary-care-contacts​/general-medical-services​/quality-and-outcomes-framework​/changes-to-qof-2012-13, accessed 13 March 2017.
  • NHS England Patient Safety Domain. Revised Never Events Policy and Framework. 2015.
  • OECD. OECD Health Systems Characteristics Survey: Section 10: Pay-for-performance and other financial incentives for providers. 2016. Available at: https://qdd​.oecd.org/subject​.aspx?Subject=hsc.
  • Ogundeji YK, Bland JM, Sheldon TA. The effectiveness of payment for performance in health care: a meta-analysis and exploration of variation in outcomes. Health Policy. 2016;120(10):1141–50. [PubMed: 27640342]
  • Olsen CB, Brandborg G. Quality Based Financing in Norway: Country Background Note: Norway. Norwegian Directorate of Health; 2016.
  • Petersen LA, et al. Does Pay-for-Performance Improve the Quality of Health Care. Annals of Internal Medicine. 2006;145(4):265. [PubMed: 16908917]
  • Rashidian A, et al. Pharmaceutical policies: effects of financial incentives for prescribers. Cochrane Database of Systematic Reviews. 2015;(8):CD006731. [PMC free article: PMC7390265] [PubMed: 26239041]
  • Robinson JC. Theory and Practice in the Design of Physician Payment Incentives. Milbank Quarterly. 2001;79(2):149. [PMC free article: PMC2751195] [PubMed: 11439463]
  • Roland M, Dudley RA. How Financial and Reputational Incentives Can Be Used to Improve Medical Care. Health Services Research. 2015;50 Suppl 2:2090. [PMC free article: PMC5338201] [PubMed: 26573887]
  • Roland M, Guthrie B. Quality and Outcomes Framework: what have we learnt. BMJ (Clinical Research edition). 2016;354:i4060. [PMC free article: PMC4975019] [PubMed: 27492602]
  • Rosenthal MB, et al. Paying For Quality: Providers’ Incentives For Quality Improvement. Health Affairs. 2004;23(2):127. [PubMed: 15046137]
  • Roski J, et al. The impact of financial incentives and a patient registry on preventive care quality: increasing provider adherence to evidence-based smoking cessation practice guidelines? Surveys available upon request from corresponding author. Preventive Medicine. 2003;36(3):291. [PubMed: 12634020]
  • Rynes SL, Gerhart B, Parks L. Personnel psychology: performance evaluation and pay for performance. Annual Review of Psychology. 2005;56:571. [PubMed: 15709947]
  • Sabatino SA, et al. Interventions to increase recommendation and delivery of screening for breast, cervical, and colorectal cancers by healthcare providers systematic reviews of provider assessment and feedback and provider incentives. American Journal of Preventive Medicine. 2008;35(1) Suppl:S67–74. [PubMed: 18541190]
  • Salize HJ, et al. Cost-effective primary care-based strategies to improve smoking cessation: more value for money. Archives of Internal Medicine. 2009;169(3):230–5. [discussion 235–6] [PubMed: 19204212]
  • Scott A, et al. The effect of financial incentives on the quality of health care provided by primary care physicians. Cochrane Database of Systematic Reviews. 2011;(9):CD008451. [PubMed: 21901722]
  • Sorbero ME, et al. Assessment of Pay-for-Performance Options for Medicare Physician Services: Final Report. Washington, DC: RAND Corporation; 2006.
  • Srivastava D, Mueller M, Hewlett E. Better Ways to Pay for Health Care. Paris: OECD Publishing; 2016.
  • Town R, et al. Economic incentives and physicians’ delivery of preventive care: a systematic review. American Journal of Preventive Medicine. 2005;28(2):234. [PubMed: 15710282]
  • van Herck P, et al. Systematic review: effects, design choices, and context of pay-for-performance in health care. BMC Health Services Research. 2010;10:247. [PMC free article: PMC2936378] [PubMed: 20731816]
  • Walker S, et al. Value for money and the Quality and Outcomes Framework in primary care in the UK NHS. British Journal of General Practice. 2010;60(574):e213–20. [PMC free article: PMC2858553] [PubMed: 20423576]
  • Wharam JF, et al. High quality care and ethical pay-for-performance: a Society of General Internal Medicine policy analysis. Journal of General Internal Medicine. 2009;24(7):854. [PMC free article: PMC2695523] [PubMed: 19294471]

Footnotes

1

The review by Kondo et al. (2015) evaluated effects of both primary and hospital care but the presentation of the results was split between Table 14.3 and Table 14.5.

© World Health Organization (acting as the host organization for, and secretariat of, the European Observatory on Health Systems and Policies) and OECD (2019)
Bookshelf ID: NBK549278

Views

  • PubReader
  • Print View
  • Cite this Page
  • PDF version of this title (4.8M)

Other titles in this collection

Related information

  • PMC
    PubMed Central citations
  • PubMed
    Links to PubMed

Recent Activity

Your browsing activity is empty.

Activity recording is turned off.

Turn recording back on

See more...