A Guide to Critically Assessing Information (part 2)

HomeInsightsBlogs | Last Updated August 5, 2021 - by brenda hillebrandt under data science & analytics

Published onJune 30, 2021

Now that you have read how to determine which internet sources are reliable in “A Guide to Critically Assessing Information (part 1),” we dive deeper into investigating the information available within those sites. Even the most well-intentioned sources can have data that may not hold up to academic scrutiny. We at Softcrylic understand the importance of data integrity and are held to the highest standards of data collection. We share these best practices with you so that you can critically assess sources on your own data discovery journey. Read on to learn how to examine the wealth of information that the internet provides.

Phrasing the Survey Questions

When designing a survey, the nuances of the questions can sway the outcome of the entire study. “It’s not what you say; it’s how you say it.” Or, in this case, “It’s not what you want to know, but how you ask it.” Pretend that an imaginary company, Marketing, Inc., is choosing a background color for a client’s website. They ask customers a very simple question, “Do you like the color blue?” 85% of the people do not dislike blue, so they respond “yes.” Now imagine that they asked, “Do you prefer red or blue?” 51% preferred red. Based on these answers, Marketing, Inc. may now wish to change the entire site color scheme to red, even though 85% of customers would have been happy with blue. These different questions give us different information, so we must be vigilant when interpreting results.

Survey Characteristics

Once you have determined that the questions were asked in an unbiased manner, you must ask yourself if the results of the survey would be applicable to your research topic.

First, was the survey conducted in a reasonable time frame? Studies performed months, weeks, or even hours in the past could be obsolete by the time you are reading the results. In this day and age, information moves at light speed, and we must move quickly to keep up. For instance, if Marketing, Inc. has a client whose website sells clothing, studies from the winter season may be useless when determining which items should go on sale for summer. “Do you prefer a fur-lined hood?” is not an applicable piece of information during swimsuit season. To accommodate annual, weekly, or even hourly trends, data scientists must identify and account for these cyclical variations when reporting data.

Second, was the survey conducted with an appropriate audience?  Following the same example, interviewing Florida residents about the best winter coat may not provide very useful information. Income brackets, geography, lifestyles, age, gender, and countless other demographic markers often segment survey respondents. Data scientists can use these to their advantage to provide clients with the most pertinent information.

Sample Size (n)

Next, it is imperative to examine sample size when determining if survey data is reliable. In a simple example, if Marketing, Inc. claims that 100% of their clients had increased email opens after one month of working together, we need to know how many total clients they have. Finding out that Marketing, Inc. only has two clients makes the claim far less impressive than if they had 100 clients with increased web traffic.

Additionally, “increased opens” could mean that the business saw a 20% increase, or just one additional email open. For the increase to be meaningful, it must pass a test of statistical significance. An event is considered statistically significant if it occurs with such frequency in the data set that it is unlikely to be due to chance alone. A common tool in a data scientist’s arsenal for determining such significance is the A/B test. In this style test, two versions of something (such as an email) is deployed to 200 recipients: 100 receive the original email (version A), 100 receiving Marketing, Inc.’s version (version B). Version A of the email had 54 recipients who purchased something; Version B, 63. Did Marketing, Inc. cause an increase in sales? While it is possible, it is also possible that this was simply due to chance. Using a statistical significance calculator, it is determined that we are 90% confident that this did not occur simply by chance. However, it does not reach the 95% confidence level threshold, a standard across most industries. By increasing the sample size to 2,000 (from 200), if the same proportion of results were achieved (540 conversions from version A and 630 from version B), we can conclude with 99% confidence that this outcome was not due to chance. It is statistically probable that, all else constant, Marketing, Inc.’s changes caused an increase in conversion rate, which we can confidently say when using the larger sample size. The math is messy and will not be discussed here, but there are many online tools for calculating statistical significance. The more confident you can be that the results are not due to chance, the more convincing the argument being made.

*The definition of “sufficiently large” varies depending on topic

Sampling Bias

While large sample sizes are usually better, not all samples are created equal. There are several phenomena that can cloud your conclusions known as “sampling biases.” Stay tuned for the next article, “A Guide to Critically Assessing Information (part 3),” in this series to learn about sampling biases.

Conclusion

Softcrylic collects millions of data points each day—web traffic, sentiment analyses, satisfaction surveys—allowing us to gather important insights that our clients use to better their businesses. To properly set up this process, we must consider all the above guidelines, along with countless others, to ensure the integrity and usefulness of the information we acquire. Asking the right questions, asking the right people, and asking the right number of them helps us Make Data Work. These best practices are applicable in many situations with data of all types, so get out there and Make Data Work for you too!

Brenda Hillebrandt

Brenda Hillebrandt is an excellent Data Analyst on our Data Science & Advanced Analytics team. Her applied knowledge of data and economics makes her great with large amounts of our client data and analyzing their data analytics.

Contact Us

We're not around right now. But you can send us an email and we'll get back to you, asap.

Not readable? Change text. captcha txt

Start typing and press Enter to search