Decision Mechanics

Insight. Applied.

  • Services
    • Decision analysis
    • Big data analysis
    • Software development
  • Articles
  • Blog
  • Privacy
  • Hire us

Survival bias

November 9, 2016 By editor

This strip from Saturday Morning Breakfast Cereal is the best description of survival bias I’ve come across.

It’s a pernicious problem in day-to-day decision-making. Modern, sensationalist news reporting reinforces it. Frightening, surprising. “man bites dog” stories survive. Stories about everyday dangers end up in the bin.

When I was a child, my parents would tell me to go to school through fields and woodland so I was away from traffic as I dawdled around. Nowadays parents want their children in busy public areas when alone. Unless child attackers are now more numerous than careless drivers, this isn’t good advice.

Filed Under: Decision science Tagged With: survival bias

RIP polling

November 9, 2016 By editor

voting intention survey

Polling died last night.

It’s been terminally ill for a while now. The predictions for the UK general election in 2015 were abysmal. Brexit polls were unreliable. And polls put the Scottish independence referendum result in 2014 as a close call when it was a resounding “No”.

After its performance in last night’s US presidential election, we have to reach for the life support switch.

Here are the poll results immediately prior to the election.

PollWinnerMargin
Monmouth UniversityClinton+6
Lucid/The Times-PicayuneClinton+5
ABC News/Washington PostClinton+4
Fox NewsClinton+4
Insights WestClinton+4
New York Times/CBS NewsClinton+4
YouGov/EconomistClinton+4
Bloomberg/SelzerClinton+3
RasmussenClinton+2
IBD/TIPPTrump+2

FiveThirtyEight, which became famous for its predictions in previous US elections, had Clinton with a 71.4% chance of winning over Trump’s 28.6%. It also had Clinton comfortably winning the popular vote by 48.5% against 44.9%.

The Independent reported on a model from Moody’s Analytics that has correctly picked every president since 1980. Well…had. It forecast Clinton would pick up 332 Electoral College seats to Trump’s 206.

Sam Wang of the Princeton Election Consortium called the result on 18 October—a win for Clinton. He tweeted

It is totally over. If Trump wins more than 240 electoral votes, I will eat a bug.

That’s lunch today sorted then.

So much time, energy and column inches have been spent on techniques which, again and again, have come up short. We can still have these polls, but we need to get them off the front page and make space for them under the horoscopes.

The problem is that we all desperately want to know what is going to happen. So, there’s a market for people who profess to be able to tell us. Given the demand, if we are going to euthanize polling, we need a replacement.

Prediction markets seemed promising. However, when I looked at Betfair Predicts before the election it was giving Clinton an 83% chance of success. An average of nine predication markets (including Betfair) published just before the election gave Clinton a 82.5% chance of victory.

So much for that then.

Polling (and betting) is based on obtaining people’s opinions—ideally a lot of people’s opinions. Unfortunately, when we lack any reasonable precedent for a situation or decision, it’s very hard to have any kind of informed opinion. Becoming informed about complex socio-economic situations takes resources—an investment very few are willing to make just to enhance the accuracy of a one-off prediction of an event they can’t change.

Crowds have no wisdom when the individuals don’t have a clue.

What about those who put a bit more time into their opinions? Well, the superforcasters at the Good Judgement Project reported a 76% chance that a Democrat would win and a 64% chance that the Democrats would control the Senate.

So much for that then.

Is there anything we can do to predict the outcomes of these elections?

As physicist and Nobel laureate Neil Bohr said

Prediction is very difficult, especially about the future.

If we can’t rely on judgement the only way forward would seem to be to improve our ability to model and study the social physics of complex systems—such as the research published in the Journal of Artificial Societies and Simulation. We’re currently a long way from being able to use such approaches with any degree of confidence, but the techniques used in this field, such as agent-based simulation, have the potential to make predictions in novel situations.

Such techniques also tend to be expensive to use—especially when compared with running an online survey. However, we can’t just keep doing things that clearly aren’t working just because we can.

RIP polling. I won’t mourn you.

Filed Under: Decision science Tagged With: forecasting, polling, prediction, US presidential election

Visualizing election results

November 4, 2016 By editor

US flag in the shape of a map of the US

The New York Times published an article today on how it has mapped election results over the years. It illustrates the challenges of trying to present complex information succinctly to a lay, and possibly hostile, audience.

As they note, simply shading the states of the US based on the party that it voted for makes the country look decidedly Republican.

A timely reminder, if we needed it, of the challenges data scientists face in their goal of providing objective, unbiased information.

Flag image by DrRandomFactor.

Filed Under: Data science Tagged With: election results, US presidential election, visualization

Sharing R code using R-Fiddle

November 3, 2016 By editor

If you want to share a snippet of R code with others—e.g. for teaching or to get help on Stack Overflow—consider using R-Fiddle.

While gists are good for basic code sharing, R-Fiddle allows others to execute the code in place. You can even embed the code together with a working R console in blog posts (as an iframe).

Filed Under: Data science Tagged With: R, R-Fiddle

RStudio 1.0 released

November 2, 2016 By editor

RStudio have released version 1.0 of their eponymous R IDE. They are calling it their

…biggest [release] ever!

It certainly has a number of very significant features.

Integrated support for Spark

Spark and R are core tools for data scientists. While Spark has an R API, support for the machine learning libraries is lagging.

So, it’s great to hear that RStudio now has integrated support for Spark and the sparklyr package. sparklyr provides extensive access to Spark’s Machine Learning Library (MLlib) and, through the rsparkling extension package, access to H2O’s distributed machine learning algorithms.

RStudio can be used to manage connections to Spark and run R functions on data held in the cluster. Data is read and transformed using Hadley Wickham’s excellent dplyr data manipulation package.

R Notebooks

R Notebooks allow the creation of documents where computation can be interspersed with narrative. Code can be executed interactively and the document updated accordingly. Readers of an R Notebook can modify the code in-place, execute it and see the new output—e.g. an updated chart. This is a particularly powerful tool for teaching R and data science.

Code profiling

I’ve used the profvis package many times to rescue clients from an analysis tool that takes hours to run. profviz provides an interactive graphical display of where you R code is spending time or eating memory.

This has now been integrated into RStudio, so you can select a block of code, click a menu option and see a visual representation of your code’s performance characteristics.

What are you waiting for?

RStudio 1.0 is free and available now on Linux, OS X and Windows. Why are you still reading this? Go and download it.

Filed Under: Data analysis, Machine learning Tagged With: R, R Notebooks, RStudio, Spark, sparklyr

  • « Previous Page
  • 1
  • …
  • 22
  • 23
  • 24
  • 25
  • 26
  • …
  • 59
  • Next Page »

Copyright © 2026 · Decision Mechanics Limited · info@decisionmechanics.com