Responsibilities
- Create data processing workflows and machine learning systems tailored for extensive NLP and large language model use cases
- Build and manage pipelines to handle and evaluate large volumes of data
- Examine complex time series data to generate practical insights and technical solutions
- Develop, deploy, and sustain models and algorithms driven by data
- Invent methods for curating datasets to enhance the effectiveness and quality of AI training
- Construct scalable systems to filter, deduplicate, and refine large text corpora for LLM training
- Work with interdisciplinary teams to identify data requirements and deliver accurate, timely results
- Maintain high standards of data accuracy and consistency across all stages
- Apply OCR technology to transform diverse document formats into usable, searchable data
- Stay current with advancements in data science and machine learning, integrating proven techniques
- Support the evolution of internal tools and infrastructure for data processing
Benefits
- competitive salary
- benefits
- equity package
Compensation
competitive salary
Work Arrangement
fully remote
Team
remote team
Other
fully remote team