Decision Mechanics

Insight. Applied.

  • Services
    • Decision analysis
    • Big data analysis
    • Software development
  • Articles
  • Blog
  • Privacy
  • Hire us

RStudio 1.0 released

November 2, 2016 By editor

RStudio have released version 1.0 of their eponymous R IDE. They are calling it their

…biggest [release] ever!

It certainly has a number of very significant features.

Integrated support for Spark

Spark and R are core tools for data scientists. While Spark has an R API, support for the machine learning libraries is lagging.

So, it’s great to hear that RStudio now has integrated support for Spark and the sparklyr package. sparklyr provides extensive access to Spark’s Machine Learning Library (MLlib) and, through the rsparkling extension package, access to H2O’s distributed machine learning algorithms.

RStudio can be used to manage connections to Spark and run R functions on data held in the cluster. Data is read and transformed using Hadley Wickham’s excellent dplyr data manipulation package.

R Notebooks

R Notebooks allow the creation of documents where computation can be interspersed with narrative. Code can be executed interactively and the document updated accordingly. Readers of an R Notebook can modify the code in-place, execute it and see the new output—e.g. an updated chart. This is a particularly powerful tool for teaching R and data science.

Code profiling

I’ve used the profvis package many times to rescue clients from an analysis tool that takes hours to run. profviz provides an interactive graphical display of where you R code is spending time or eating memory.

This has now been integrated into RStudio, so you can select a block of code, click a menu option and see a visual representation of your code’s performance characteristics.

What are you waiting for?

RStudio 1.0 is free and available now on Linux, OS X and Windows. Why are you still reading this? Go and download it.

Filed Under: Data analysis, Machine learning Tagged With: R, R Notebooks, RStudio, Spark, sparklyr

RDDs, DataFrames and Datasets

July 14, 2016 By editor

There are now three Spark APIs for working with large volumes of data

  • RDD
  • DataFrame
  • Dataset

Which one should we use? Good question. Jules Damji provides a pretty comprehensive answer in an article on the Databricks blog.

RDD was the original API for working with large volumes of data. The first thing to note is that the RDD API is not being deprecated. It has an important role to play. RDDs make sense when working with unstructured data, such as media or text streams. They are also the best approach if your problem fits neatly within the functional programming paradigm.

However, for the majority of data science tasks, it is likely that the DataFrame and Dataset APIs will be more appropriate. Dataset is a strongly-typed API, whereas DataFrame is untyped. A DataFrame can be thought of as a Dataset of generic (untyped) objects. From Spark 2.0 onward the Dataset and DataFrame APIs will be unified.

Datasets imposes more constraints on the structure of the data. They are not as flexible as RDDs. However, those constraints allow the API to have higher-level functionality and support enhanced compile-time checks and significant run-time performance optimizations.

So, at the risk of oversimplifying, use the Dataset API unless it’s making you jump through hoops. If it is, feel free to use the RDD API. It’s not disappearing anytime soon.

It should be noted that the Spark libraries (such as MLlib) are still being updated to work with the Dataset API, so, in the short term, RDDs may still make sense even when working with structured data.

Filed Under: Big data, Data science, Machine learning Tagged With: DataFrame, dataset, RDD, Spark

Status of Spark MLlib wrappers in SparkR

July 14, 2016 By editor

Wrappers for Spark’s MLlib machine learning library in SparkR have been slow to arrive. However, the future looks bright.

The imminent 2.0 release will bring k-means support to SparkR and the 2.1 release is scheduled to include wrappers for the following machine learning stalwarts

  • Alternating Least Squares (ALS)
  • Decision Trees
  • Gaussian Mixture Models
  • Isotonic Regression
  • Latent Dirichlet Allocation (LDA)
  • Multilayer Perceptron Classifiers
  • Random Forests

Filed Under: Machine learning Tagged With: mllib, Spark, sparkr

Free Apache Spark Analytics Made Simple e-book

March 31, 2016 By editor

Apache Spark Analytics Made Simple-e-book cover

Databricks have just published a free e-book entitled “Apache Spark Analytics Made Simple”. Contents include

  • An introduction to the Spark API for analytics
  • Tips and tricks to simplify unified data access
  • Real-world case studies of how various companies are using Spark with Databricks to transform their business

There are more to come. Titles are

  • Mastering Advanced Analytics with Apache Spark
  • Lessons for Large Scale Machine Learning Deployments on Apache Spark
  • Building Real-Time Applications with Spark Streaming

Filed Under: Big data, Data analysis Tagged With: e-book, Spark

Copyright © 2026 · Decision Mechanics Limited · info@decisionmechanics.com