Using internet search data to predict new HIV diagnoses in China: a modelling study

BMJ Open. 2018 Oct 17;8(10):e018335. doi: 10.1136/bmjopen-2017-018335.

Abstract

Objectives: Internet data are important sources of abundant information regarding HIV epidemics and risk factors. A number of case studies found an association between internet searches and outbreaks of infectious diseases, including HIV. In this research, we examined the feasibility of using search query data to predict the number of new HIV diagnoses in China.

Design: We identified a set of search queries that are associated with new HIV diagnoses in China. We developed statistical models (negative binomial generalised linear model and its Bayesian variants) to estimate the number of new HIV diagnoses by using data of search queries (Baidu) and official statistics (for the entire country and for Guangdong province) for 7 years (2010 to 2016).

Results: Search query data were positively associated with the number of new HIV diagnoses in China and in Guangdong province. Experiments demonstrated that incorporating search query data could improve the prediction performance in nowcasting and forecasting tasks.

Conclusions: Baidu data can be used to predict the number of new HIV diagnoses in China up to the province level. This study demonstrates the feasibility of using search query data to predict new HIV diagnoses. Results could potentially facilitate timely evidence-based decision making and complement conventional programmes for HIV prevention.

Keywords: health informatics; internet; predictive model; search query; surveillance.

Publication types

  • Research Support, N.I.H., Extramural
  • Research Support, Non-U.S. Gov't

MeSH terms

  • Bayes Theorem
  • China / epidemiology
  • Forecasting*
  • HIV Infections / diagnosis
  • HIV Infections / epidemiology*
  • Humans
  • Internet*
  • Models, Statistical*
  • Prevalence
  • Search Engine