Linguist II (AI/LLM)
Prolim global system · Remote · Baltimore
Job description
We are seeking a Linguist II to support the development of AI-powered products, including Large Language Models (LLMs) and voice-enabled technologies. This role is ideal for candidates with a strong background in linguistics, language technologies, NLP, and AI data operations who are passionate about improving language understanding and model performance.
The successful candidate will collaborate with cross-functional teams to develop, evaluate, and improve multilingual datasets used for AI model training, fine-tuning, and evaluation.
Key Responsibilities
- Apply linguistic expertise in syntax, semantics, pragmatics, sociolinguistics, and corpus linguistics to support AI and LLM development.
- Collaborate with linguists, data operations teams, and machine learning engineers on data collection, annotation, localization, and curation.
- Create and maintain annotation guidelines and schemas for AI training datasets.
- Evaluate and quality-check datasets used for model training, fine-tuning, and alignment.
- Generate and validate synthetic annotated data to improve AI model performance.
- Perform prompt testing, linguistic error analysis, red-teaming, and model evaluation.
- Review AI-generated content for hallucinations, bias, tone, and linguistic quality.
- Support Responsible AI initiatives through linguistic analysis and quality assurance.
- Participate in experiments to measure annotation quality, data consistency, and downstream model performance.
Required Qualifications
- Bachelor's degree in Linguistics, Computational Linguistics, Computer Science, Speech Science, or a related field.
- 1+ years of experience in Linguistics, Language Technologies, NLP, AI/ML data operations, or a related field.
- Native or near-native English proficiency plus fluency in at least one additional language.
- Strong understanding of:
- Syntax
- Semantics
- Pragmatics
- Sociolinguistics
- Corpus Linguistics
- Familiarity with Large Language Models (LLMs), prompting, evaluation, training data, and fine-tuning.
- Experience with semantic ontologies, taxonomies, or intent/slot frameworks.
- Experience using AI assistants, chatbots, or AI agents.
- Knowledge of SQL, spreadsheets, R, Unix, or other data analysis tools.
- Experience working with multilingual speech and text datasets.
- Excellent analytical, communication, and collaboration skills.
Preferred Qualifications
- Master's degree in Linguistics, Computational Linguistics, Language Technologies, or a related discipline.
- Experience with NLP libraries and machine learning tools such as:
- Hugging Face
- spaCy
- NLTK
- PyTorch
- Familiarity with statistical language modeling and AI training data pipelines.
- Strong organizational skills with exceptional attention to detail.
Preferred Technical Skills
- Large Language Models (LLMs)
- Natural Language Processing (NLP)
- AI Data Annotation
- Prompt Engineering
- Model Evaluation
- SQL
- Python (preferred)
- Hugging Face
- spaCy
- NLTK
- PyTorch
- AI Agents / Chatbots
- Data Quality & Annotation
Pay: $50.00 - $60.00 per hour
Work Location: Remote
ML/AI Work links you to the employer's original posting — always verify the details there before applying.
More Core AI Engineering roles
View all →GenAI / Agentic AI Engineer (US)
TD · Philadelphia, US
Senior AI Engineer, Scientific Training & Collaboration
— · San Jose, US
AI Engineer, Government Solutions & APIs
— · San Jose, US
AI Solutions Engineer - Defense Tech (Secret Clearance)
— · Baltimore, US
AI Solutions Engineer
Innodata · Baltimore, US
AI Developer
ARCHE consulting · Remote · Poznań