About the job
As a Data Scientist on our team, you will analyze data from massive data sets to categorize customer idiosyncrasies, identify outliers, and systematically detect anomalies that substantially affect the performance of our models. You will work closely with other senior technical leaders within the team and across AWS. You should know how to trace decisions in data from raw data through complex models to their impact business metrics. Experience with machine learning explainability is a plus. You should be able to translate well-defined business problems into data science problems and you solve these problems using appropriate assumptions, methodologies, and data science best practices. Our team is at an early stage, so you will have significant impact on our deliverables with no operational load from existing models/systems.
Responsibilities
Categorizing Customer Idiosyncrasies: As we expand to more customers, we are discovering that they use our product in very different ways and that poses issues for our models. Effectively summarizing these differences (for example, X% of customers do Y) would be immensely helpful.
Detecting and Cleaning Up Outliers: We have situations where outliers have a huge impact on model outputs. You will help us develop mechanisms to clean up outliers for downstream consumption.
Deep Diving Customer Issues: Customers have longstanding traditions and trusted formulas for managing their contact centers. When our formulas differ from theirs, we need to deep dive these discrepancies and determine if there is an issue with our model or if we are giving the customer better results than they are used to.
Assessing Data Gaps: It's hard to estimate the weather in Seattle if the only data you have is the average weight of elephants in Zimbabwe. We know we don't have all the data we need, but we need to answer two related questions: (a) what features can we derive in creative ways from existing data sources? (b) can we estimate the benefit of getting a new data stream in terms of accuracy improvement?
Qualifications
Minimum
- 2+ years of data scientist experience
- 3+ years of data querying languages (e.g. SQL), scripting languages (e.g. Python) or statistical/mathematical software (e.g. R, SAS, Matlab, etc.) experience
- 3+ years of machine learning/statistical modeling data analysis tools and techniques, and parameters that affect their performance experience
- Experience applying theoretical models in an applied environment
Preferred
- Experience in Python, Perl, or another scripting language
- Experience in a ML or data scientist role with a large technology company