Top 3 Sessions I’m Excited for at RStudio Conference 2020
On January 29th – January 30th RStudio will be holding their annual conference. Last year the conference was in Austin, Texas. However, this year the conference will find itself in the archetypal technology hub of the United States – San Francisco, California. SF is a fitting choice for the R language’s preeminent open source software company, as San Francisco is almost ubiquitously considered the top technology hub in the United States, if not the world.
Despite the location change, this year’s conference still boasts the same excellent ensemble of speakers and sessions. If you’d like to compare this year’s lineup to last year’s, be sure to check out the RStudio Conference 2019’s video archive, where you will find all kinds of goodies. Now let’s dive into the sessions I am most excited about.
Totally Tidy Tuning Techniques | Max Kuhn
Max Kuhn was the creator of the caret package and is currently the architect of the tidymodels machine learning ecosystem at RStudio. Anyone who is familiar with me knows I’m a huge fan of Max Kuhn. His book Applied Predictive Modeling was incredibly influential in my development as a Data Scientist and it is always my first recommendation for anyone interested in the field. Be sure to check out his new book Feature Engineering and Selection: A Practical Approach for Predictive Models!
In this session, Max will be going into two packages within the tidymodels meta package, tune & workflows. Specifically, how to use these packages to efficiently choose hyperparameters via grid search or Bayesian optimization.
Can’t wait!
We’re hitting R a million times a day so we made a talk about it | Heather Nolis & Dr. Jacqueline Nolis
The duo speaking again about their use of R in production at T-Mobile following over a year of having their model deployed, based on their excellent series of medium posts (Part 1, Part 2, Part 3) and talk at last year’s conference on the topic. Their series – which ought to be required reading for anyone interested in deploying R in production – highlights the steps and services their team leveraged to deploy their model, easing the reader into the world of REST APIs, Docker, and Amazon Web Services.
Accelerating Analytics with Apache Arrow | Neal Richardson
Apache Arrow is a project that’s focused on creating a universal, language-independent, columnar memory format which can be shared between different programming languages and processing engines. Ideally, Arrow aims to standardize file formatting across the data science and engineering community, eliminating some of the overhead caused by serialization.
The Arrow format has already shown evidence to support its efficacy towards the objective of providing a unified format. One example of this is when Arrow’s being used with pandas to pass a python User Defined Function (UDF) to transform data in spark. Under normal conditions, passing a python UDF to Spark is a last resort because of the immense serialization cost, but by leveraging arrow this overhead is essentially eliminated (Slide 16).
The promising preliminary results and the great reputation of the team behind the Apache Arrow project have me incredibly excited to hear directly from someone actively involved with the project.
Feel free to come up to Brad Kossmann and me to chat if you are attending! Safe Travels!