Epidemic modelling with compartmental models using R

Posted on December 11, 2012 by admin

[After reading through this module you should have an intuitive understanding of how infectious disease spreads in the population, and how that process can be described using a compartmental model with flow between the compartments. You should be able to write down the differential equations of a simple disease model, and you will learn in this module how to numerically solve those differential equations in R to obtain the model estimate of the epidemic curve]

An excellent reference book with background material related to these lectures is Mathematical Epidemiology by Brauer et al.

Contents:

Introduction
Basic dynamics of infectious disease spread
The SIR compartmental model of disease spread
The SIR model system of equations
Numerically solving the SIR model system of equations in R
R code to model an influenza pandemic with an SIR model
Further things you can explore
Summary

Introduction

Models of disease spread can yield insights into the mechanisms and dynamics most important to the spread of disease (especially when the models are compared to epidemic data). With this improved understanding, more effective disease intervention strategies can potentially be developed. Sometimes disease models are also used to forecast the course of an epidemic, and doing exactly that for the 2009 pandemic was my introduction to the field of computational epidemiology.

There are lots of different ways to model epidemics, and there are several modules on this site on the topic, but let’s begin with one of the simplest epidemic models for an infectious disease like influenza: the Susceptible, Infected, Recovered (SIR) model.

Continue reading →

The basics of the R statistical progamming language

Posted on December 11, 2012 by admin

[After you have read through this module, and have downloaded and worked through the provided R examples, you should be proficient enough in R to be able to download and run other R scripts that will be provided in other posts on this site. You should understand the basics of good programming practices (in any language, not just R). You will also have learned how to read data in a file into a table in R, and produce a plot.]

Contents:

Why use R for modelling?
How to download R
Some example R code with an overview of basic R commands
Advancing on: programming constructs
Good programming practices (in any language)
Reading data files into R

Why use R for modelling?

I have programmed in many different computing and scripting languages, but the ones I most commonly use on a day to day basis are C++, Fortran, Perl, and R (with some Python, Java, and Ruby on the side). In particular, I use R every day because it is not only a programming language, but also has graphics and a very large suite of statistical tools. Connecting models to data is a process that requires statistical tools, and R provides those tools, plus a lot more.

Unlike SAS, Stata, SPSS, and Matlab, R is free and open source (it is hard to beat a package that is more comprehensive than pretty much any other product out there and is free!).

Continue reading →

Welcome to Polymatheia

Posted on December 11, 2012 by admin

I am a data scientist with a diverse background visual analytics, data mining, social media analytics, machine learning, high performance computing, and mathematical and computational dynamical modeling. As both an academic and a consultant, I have worked on a broad array of research topics in public health and the social sciences, including crime and violence risk analyses, spread of political and partisan sentiments in a society, disease modelling, and other topics in applied modelling in the social and life sciences, with over 360 publications on a wide variety of subjects. My unique trans-disciplinary skill set enables me to examine a wide range of research questions that are often of broad interest and importance to policy makers and the general public.

My work in the computational social sciences has been high impact, including an analysis of how media can incite panic in a population, and how contagion may play a role in the temporal patterns observed in mass killings in the US. The latter work has received extensive media coverage, including on BBC, CNN, NBC, CBS, FOX, NPR, Reuters, USA Today, Washington Post, Los Angeles Times, The Wall Street Journal, Huffington Post, Christian Science Monitor, Newsweek, The Atlantic, The Telegraph, The Guardian, Vox, and many other local, national, and international news agencies. The study was also profiled in Season 8, Episode 4 of Morgan Freeman’s Through the Wormhole documentary series on the Discovery Channel in 2017, “Is Gun Crime a Virus?”.

Several of my publications in computational epidemiology have also received widespread attention, including forecasts for the progression of the 2009 influenza pandemic, and the 2014 Ebola epidemic in West Africa. Both analyses correctly forecast the progression of the epidemics, and my forecast for the 2009 pandemic was discussed in a special US Senate meeting on October 23, 2009 H1N1 Flu: Monitoring the Nation’s response.

Through my consulting company, Towers Consulting LLC, I provide consulting services for industry, academia, and the public sector in quantitative and predictive analytics, modelling, statistics, and visual analytics. Notable past and current clients include the VACCINE Department of Homeland Security Center of Excellence, Cure Violence Global, the Carter Center, George Washington University, and Disney. Prospective clients can contact me at TowersConsultingLLC@gmail.com.

On this website I share information and computational tools related to a wide range of my research topics, including visual analytics, applied statistics, and mathematical and computational modelling methods. This website also includes material related to my lectures and seminars in statistical and computational methods in the life and social sciences.

My CV is available here.

Finding sources of data for computational, mathematical, or statistical modeling studies: free online data

Posted on April 3, 2012 by admin

[In this module we discuss methods for finding free sources of online data. We present examples of climate, population, and socio-economic data from a variety of online sources. Other sources of potentially useful data are also discussed. The data sources described here are by no means an exhaustive list of free online data that might be useful to use in a computational, statistical, or mathematical modeling study.] Continue reading →

Finding sources of data: extracting data from the published literature

Posted on April 2, 2012 by admin

Connecting mathematical models to predicting reality usually involves comparing your model to data, and finding model parameters that make the model most closely match observations in data. And of course statistical models are wholly developed using sources of data.

Becoming adept at finding sources of data relevant to a model you are studying is a learned skill, but unfortunately one that isn’t taught in any textbook!

One thing to keep in mind is that any data that appears in a journal publication is fair game to use, even if it appears in graphical format only. If the data is in graphical format, there are free programs, such as DataThief, that can be used to extract the data into a numerical file.

Continue reading →

Literature searches with Google Scholar

Posted on April 1, 2012 by admin

Google Scholar is search engine that indexes the scholarly literature across an array of publishing formats and disciplines. It provides a very powerful means to find literature associated with pretty much any research topic you can think of.

Continue reading →

AML 610 Fall 2014: List of modules

Posted on March 29, 2012 by admin

The syllabus for this course can be found here.

The final write-ups for final group projects are due Monday, December 1st, 2014. On Dec 2nd and 3rd students will meet with Prof Towers to receive feedback on their project and writeup.

Each of the project groups will perform an in-class 20 min presentation on Monday, Dec 8th, 2014 and Wed, Dec 10th, 2014. By Dec 9th, all group members are to submit to Prof Towers a confidential email, detailing their contribution to the group project, and detailing the contributions of the other group members.

The list of modules for the Fall 2014 course in computational and statistical methods for mathematical biologists and epidemiologists:

Literature searches with Google Scholar
Extracting data from graphs in published literature
Online sources of free data
- Homework #1, Due Sep 2nd 2014 at noon.
Module I: The basics of the R statistical progamming language
- Homework #2, Due Sep 10th 2014 at noon.
- as part of the homework, read the modules “How to write a good scientific paper”, and “How to download an R script from the internet and run it”.
Module II: Epidemic modelling with compartmental models
- Homework #3, Due Sep 17th 2014 at noon.
Module III: SIR disease model with age classes
- Homework#4, Due Sep 24th 2014 at noon.
Module IV: SIR modelling of influenza with a periodic transmission rate
Module V: fitting the parameters of an SIR model to influenza data using Least Squares and the Monte Carlo parameter sweep method
Module VI: an overview of goodness of fit statistics, and methods to fit parameters of mathematical models to data
Module VII: estimating parameter confidence intervals when using the Monte Carlo parameter sweep optimization method
- Homework#5, Due Wed Oct 8th 2014 at noon.
Module VIII: Basic Unix
Module IX: introduction to C++ for computational epidemiologists
- Homework #6, Due Mon Oct 20th 2014 at noon.
Module X: a C++ class to solve ordinary differential equations
- Homework #7, Due Mon Nov 3rd 2014 at noon. (assignment emailed)
- Getting started using the NSF XSEDE distributed computing system
- An example of submitting a batch job to XSEDE Stampede
Module XI: another example C++ program, fitting SIR model parameters to CDC Midwest 2007-08 B influenza data
- Homework #8, Due Mon Nov 10th 2014 at noon.
- Submitting jobs to the ASU A2C2 ASURE batch computing system
- Correcting the Pearson chi-squared statistic for over-dispersed data
- Homework #9, Due Wed Nov 19th 2014 at noon.
Module XII: another example of the parameter sampling model optimization method: using Negative Binomial likelihood
- How to ssh and scp between Unix machines without passwords
- R and C++ code related to our class publication project, and running the jobs in batch on the ASU A2C2 ASURE supercomputing system
- Homework #10, Due Mon Dec 1st 2014 at noon.
Module XIII: practical problems when connecting deterministic models to data
Module XIV: submitting jobs in batch to the ASU Saguaro distributed-computing system

AML610 Fall 2013 lecture series

Posted on August 21, 2011 by admin

Introduction

In this course students will be introduced to statistical modelling methods such as linear regression, factor regression, and time series analysis. All modelling and data analysis will be performed in the R statistical programming language. The course meets on Thursdays from 12:00-2:45 pm in PSA 546. The course syllabus can be found here.

The course will be structured in a series of modules covering various topics. Some modules may take more than one lecture to cover. Homework will be occasionally assigned throughout the course, usually after completion of a module.

The final project for the course will account for 50% of the grade, and is required to be based on a statistical analysis of one or more data sets from this page and in conjunction with discussion with myself and perhaps other faculty to determine if the topic of the analysis is novel. Students are encouraged to work together on the project in groups of up to three, but the contributions of each student in the group to the project must be clearly defined. The final project write-up will be due approximately two weeks before the end of classes (date to be announced). Oral presentations of the projects will take place in class during the last week of classes.

There is no one textbook that covers the material in this course. Since students in this course have varying backgrounds in statistics, I strongly recommend that you go to the library and take a look at the statistical texts there, and find one or two that cover linear regression and/or time series analysis that you think are at your level. For recommended texts that I think are good, Statistical Data Analysis by Cowan is a good general text for various statistical methods. A standard textbook for linear regression is Applied Linear Statistical Models by Kutner et al. Two good textbooks that cover time series analysis is Time Series Analysis with Applications in R by Cryer, and Time Series Analysis and its Applications with R Examples by Shumway. The material covered in these texts is much more expansive in scope than the material covered in this course (because they cover material that would form the basis of at least three or four different courses). All of these texts are available off of libgen.info, but note that copyright infringement is a crime… one must never do such a terrible thing.

2013 MTBI summer institute

Posted on May 7, 2011 by admin

Homework #1

Homework #2