Making Data Work – Part I
This past year our team released our first edition of the “Making Data Work Newsletter”. In the newsletter we explore hot topics in the tech and MarTech space, exciting developments at Softcrylic and other nuggets of content that hasn’t been released in other places. Initially, we release this to our newsletter subscribers, so they get the information first! Now in the new year, we want to share it with you all! In part one we dive into spotlights from the industry that you will want to be aware of, Don’t miss out on the next edition, subscribe to our newsletter here. Enjoy!
Spotlights
In our spotlights section we outline and highlight industry specs that you need to be aware of. If you have questions on the content, are dealing with similar issues or want to discuss your organization’s use case, contact our team here.
1. The Importance of Improving Data Quality for your Business
In the world of data, businesses have the ability to uncover insights that have not been explored before. The power of those insights is limitless. However, to unleash the potential of your data, you will have to do the hard work of data cleaning and quality assurance. Data quality is one of the most important components of analytics – a high level of confidence in your data is needed for insights and dashboards built on that data to be useful.
It is important to define exactly what good data quality means. The key is to establish a set of criteria. There are a number of qualities that determine whether your data is clean. Below we have outlined the most common of these.
Setting a threshold of satisfaction in each of these characteristics is needed. Our recommendation is to focus on two key questions:
How can we detect data quality problems?
Establishing checks throughout the main layers of the data pipeline
- ETL layer
- Business Logic layer
- Reporting layer
What can you do to handle these issues?
- Prioritize and correct common errors
- Prevention checks
- Full system audits
By focusing on these recommendations for improving data quality, you can reduce the risk associated with making bad decisions based on poor data and provide stakeholders with a higher level of confidence when making data-driven decisions.
2. DevOps Vs MLOps
MLOps is an extension of DevOps and sits at the intersection of ML, DevOps, and Data Engineering. It is a set of practices that aims to build, deploy, and maintain ML models efficiently and reliably in a continuous development process.
The development and deployment of a machine learning system has several important differences from the traditional DevOps process, both in terms of artifacts, stakeholders, requirements, and processes. Primary Stakeholders in MLOps are Data Scientists, DevOps Engineers, and Machine Learning engineers and are responsible in Productionizing the ML Solution.
MLOps introduces additional artifacts and processes into the traditional DevOps process like it allows you to track & monitor ML experiments. Some processes are Data Engineering, Reproducibility of models and predictions, model performance assessment, Monitoring, Scalability, Governance and regulatory compliance, and artifacts include Datasets, Models, Hyperparameters, Metrics and Workflows. It also allows you to package ML code in a reusable, reproducible form to share with other data scientists or transfer to production.
Most of the modern cloud providers provide off-the-shelf ML Platforms to enable MLOps. Popular ones among them include ‘AI Platform’ by Google Cloud Platform, ‘AzureML’ by Microsoft, and ‘SageMaker’ by Amazon Web Services. There are open-source platforms also like ‘MLFlow’ which is widely used for MLOps.
To learn more, please see our MLOps explanation here.
Best Practices
A look into new technologies or platforms in our industry that will improve your current practices.
Self Service Analytics as a Service
Standing up a self-service analytics capacity has historically been fraught with the challenges of establishing platforms and tools to serve citizen analysts across the lifecycle of organizational data. Compromises have always been required, and capacities have been limited to priority subsets of the lifecycle. The results have underdelivered on the promise of data driven ubiquity across the citizenry of the organization.
More recently, however, advances in the capabilities of Microsoft’s Power Platform have delivered a one-stop solution for end-to-end self-service capacities. Driven by a relentless cadence of enhancement since initial release in 2015, today the service aggregates an extensive suite of tools and capacities into a unified ecosystem for self-service success.
This thriving platform integrates the fundamental requirements of an enterprise self-service architecture across multiple tiers of service:
Azure Data Lake Gen 2 Storage provides data storage capacity, while compute for backend data operations is delivered by the azure hosted Power BI service. These data pipelines feed into in-memory data models at the core of the hosted platform.
Where scales of data prohibit loading entire models into memory, Azure Synapse SQL pools can be seamlessly leveraged to provide massively scalable composite models. These models serve cutting edge presentation experiences, leveraging best-in-class logical capabilities of DAX alongside the advanced interactive capabilities of the Power BI authoring tools in the cloud and on the desktop.
Recent enhancements to the platform include advanced features such as direct query on Power BI datasets, enabling multiple sets of existing semantic development to be integrated into new models. Additionally, Power BI datamarts now expose Power BI semantic development to SQL endpoints, enabling traditional SQL analytics on top of the development effort built into robust Power BI report Suites.
Together, these features mean that successive reporting deliveries will build out an increasingly inventory of self-service datasets and logic for consumption and socialization across the organization. These features are available across multiple tiers of service, ranging from per user licensed Pro tiers to dedicated Premium capacities.
As one of Microsoft’s elite gold partners, Softcrylic leads Power Platform engagements where these next generation capabilities are proving transformative for our customers. On engagements ranging from net new development to full platform migrations, Softcrylic boasts a proven record of expertise with the Power Platform. And together with our customers, we consistently deliver on this promise, empowering the citizens of the organization to expose data driven answers to the daily questions of the business.
Never miss another newsletter again, subscribe to our monthly newsletter to stay up to date on industry news, Softcrylic events, content and more! Subscribe to our newsletter here.