Decision Mechanics

Insight. Applied.

  • Services
    • Decision analysis
    • Big data analysis
    • Software development
  • Articles
  • Blog
  • Privacy
  • Hire us

DeepMind has access to data on millions of NHS patients

May 1, 2016 By editor

DeepMind, the Google-owned AI company, has an agreement with the Royal Free NHS Trust that gives it access to healthcare data on 1.6 million patients who visit three London hospitals. The agreement includes access to five years of historical data.

The company is building an application to help monitor patients with kidney disease.

While it will be interesting to see what DeepMind can produce with such comprehensive data, the National Health Service sharing identifiable healthcare information with a commercial entity is sure to raise privacy concerns.

Filed Under: Big data, Data analysis Tagged With: healthcare, NHS, privacy

Free Apache Spark Analytics Made Simple e-book

March 31, 2016 By editor

Apache Spark Analytics Made Simple-e-book cover

Databricks have just published a free e-book entitled “Apache Spark Analytics Made Simple”. Contents include

  • An introduction to the Spark API for analytics
  • Tips and tricks to simplify unified data access
  • Real-world case studies of how various companies are using Spark with Databricks to transform their business

There are more to come. Titles are

  • Mastering Advanced Analytics with Apache Spark
  • Lessons for Large Scale Machine Learning Deployments on Apache Spark
  • Building Real-Time Applications with Spark Streaming

Filed Under: Big data, Data analysis Tagged With: e-book, Spark

17000 UK male pregnancies reported in 2012

March 29, 2016 By editor

A 2012 study of National Health Service data in the UK found that there were

  • over 17000 male inpatient admissions to obstetric services
  • over 8000 male inpatient admissions to gynecology
  • nearly 20000 male inpatient admissions to midwifery

Before jumping to the conclusion that the UK is at the forefront of an exciting/disturbing evolutionary trend we should probably look for a simpler explanation—and it’s data coding errors.

Each procedure has an associated code and data entry errors resulted in men being assigned to female-only procedures. Obviously there are going to be all sorts of other errors, but it’s more difficult (and considerably less hilarious) to determine whether a patient actually had an ear infection when recorded as having a back problem.

This illustrates one of the biggest challenges for data science—garbage-in results in garbage-out. You can have the most sophisticated analysis algorithms available, but, if you are analyzing the wrong thing, you’ll draw the wrong conclusions.

Clearly it’s possible to perform statistical checks on the data—as the 2012 study illustrates. However, knowing that the data is wrong does little more than render it worthless. It could have, and should have, been checked at the point of entry. Simple logic in the data entry software could have checked the gender of the patient against a list of gender-specific codes and prevented the incorrect data from entering the system in the first place.

It’s always more efficient to fix data upstream.

Of course, gender checks would only catch some of the errors. Other techniques would be required for gender-neutral procedures. As I don’t know the data well enough, it’s difficult to come up with specific recommendations. However, some ideas might include the following.

  • Display the description of the code when it’s entered. If you are in a general practitioner’s clinic and “neurosurgery” pops up on the screen, you might catch that.
  • Allow a range of valid codes to be configured on a per terminal basis. If I’m in a gynecology department, assigning a flu code might be suspicious.
  • Learn what codes are entered at a given location/terminal. If I’ve never entered a dialysis procedure before, ask me to confirm it.
  • Alert me to very rare codes. What are the chances that I really have a patient with rabies in the UK?

One of the cheapest things we can do to improve the quality of data analysis is to improve the quality of data entry. Making basic checks in data entry systems is very, very far from rocket science.

Filed Under: Big data, Data analysis, Software

Overview of Microsoft R Server

March 11, 2016 By editor

Learning Tree just published my overview of Microsoft R Server.

Filed Under: Big data, Data analysis, Data science Tagged With: Microsoft R Server, R

Microsoft R Server is now on the Microsoft Data Science Virtual Machine

March 2, 2016 By editor

The Microsoft Data Science Virtual Machine (DSVM) now comes pre-configured with Microsoft R Server Developer Edition.

As you can scale the DSVM according to your needs, this is an easy way to get going with some heavy duty R computations.

Filed Under: Big data, Data analysis, Data science, Machine learning, Software

  • « Previous Page
  • 1
  • 2
  • 3
  • 4
  • 5
  • 6
  • …
  • 9
  • Next Page »

Copyright © 2026 · Decision Mechanics Limited · info@decisionmechanics.com