AWS Big Data Blog
Category: Database
Building a Binary Classification Model with Amazon Machine Learning and Amazon Redshift
Guy Ernest is a Solutions Architect with AWS This post builds on Guy’s earlier posts Building a Numeric Regression Model with Amazon Machine Learning and Building a Multi-Class ML Model with Amazon Machine Learning. Many decisions in life are binary, answered either Yes or No. Many business problems also have binary answers. For example: “Is […]
Test drive two big data scenarios from the ‘Building a Big Data Platform on AWS’ bootcamp
Matt Yanchyshyn is a Sr. Manager for AWS Solutions Architecture AWS offers a number of events during the year such as our annual AWS re:Invent conference, the AWS Summit series, the AWS Pop-up Loft, and a variety of roadshows. All of these provide opportunities for AWS customers to attend talks focused on big data and […]
Optimizing for Star Schemas and Interleaved Sorting on Amazon Redshift
Chris Keyser is a Solutions Architect for AWS Many organizations implement star and snowflake schema data warehouse designs and many BI tools are optimized to work with dimensions, facts, and measure groups. Customers have moved data warehouses of all types to Amazon Redshift with great success. The Amazon Redshift team has released support for interleaved […]
A Zero-Administration Amazon Redshift Database Loader
Ian Meyers is a Solutions Architecture Senior Manager with AWS With this new AWS Lambda function, it’s never been easier to get file data into Amazon Redshift. You simply push files into a variety of locations on Amazon S3 and have them automatically loaded into your Amazon Redshift clusters. Using AWS Lambda with Amazon Redshift […]
Building Multi-AZ or Multi-Region Amazon Redshift Clusters
This blog post was last reviewed July, 2022. This post explores customer options for building multi-region or multi-availability zone (AZ) clusters. By default, Amazon Redshift has excellent tools to back up your cluster via snapshot to Amazon Simple Storage Service (Amazon S3). These snapshots can be restored in any AZ in that region or transferred […]
Using Attunity CloudBeam at UMUC to Replicate Data to Amazon RDS and Amazon Redshift
Matt Yanchyshyn is a Principal Solutions Architect at AWS. Brad Helicher, Director of Cloud Business at Attunity, also contributed to this post. Attunity is an APN Big Data Competency Partner. Introduction University of Maryland University College’s mission is to provide a quality education at an affordable cost to busy professionals, mainly adults who are juggling […]
Using Amazon Redshift to Analyze Your Elastic Load Balancer Traffic Logs
Biff Gaut is a Solutions Architect with AWS Introduction With the introduction of Elastic Load Balancing (ELB) access logs, administrators have a tremendous amount of data describing all traffic through their ELB. While Amazon Elastic MapReduce (Amazon EMR) and some partner tools are excellent solutions for ongoing, extensive analysis of this traffic, they can require […]
Using AWS for Multi-instance, Multi-part Uploads
James Saull is a Principal Solutions Architect with AWS There are many advantages to using multi-part, multi-instance uploads for large files. First, the throughput is improved because you can upload parts in parallel. Amazon Simple Storage Service (Amazon S3) can store files up to 5TB, yet a single machine with a 1Gbps interface would take […]
Best Practices for Micro-Batch Loading on Amazon Redshift
NOTE: Amazon Kinesis Data Firehose is a fully managed service for delivering real-time streaming data to Amazon Redshift. For more information, please visit the Amazon Kinesis Data Firehose documentation page, “Choosing Amazon Redshift for Your Destination.” February 9, 2024: Amazon Kinesis Data Firehose has been renamed to Amazon Data Firehose. Read the AWS What’s New […]
Powering Gaming Applications with Amazon DynamoDB
Nate Wiger is Principal Gaming Solutions Architect for AWS. Dave Lang, Senior Product Manager for Amazon DynamoDB, also contributed to this article. Amazon DynamoDB is rapidly becoming the go-to database for many of the fastest-growing games in the world. Games like Fruit Ninja (from Halfbrick Studios) and Battle Camp (from PennyPop) have leveraged Amazon DynamoDB’s […]