"AI training" sounds like something that should be entirely done by machine learning engineers. It's crucial to know that in practice, humans play an essential role in data preparation, teaching AI models with examples of what good answers should be, checking their output and assisting in the improvement of their performance.
However, there's a key distinction: AI training is a broad term. Various companies use it to refer to different types of work, and not all of the jobs involved in AI training include the actual training of a model.
So, What Does AI Training Really Mean?
In simple terms, AI training involves using data, examples, feedback, and other methods to assist an AI system in learning to complete a task more effectively and improve its output.
Think about an AI model that needs to answer questions. Giving it thousands, or millions of pieces of information can help it to learn patterns in language. However, this isn't necessarily enough to teach it how to behave like people expect.
For instance, an AI could generate a grammatically correct response that doesn't address the user's query.
It may provide a convincing response but be incorrect at a factual level.
It might generate a lengthy response on a request for a brief answer.
It will likely misinterpret a specific professional or cultural situation.
Humans can be crucial in solving these issues, and also provide insights into how these issues should be addressed.
Then, the human input can become useful training and evaluation data for the AI model. AI training is not just a matter of putting information into a computer. It require individuals to make quality, accurate, relevant, safe, and useful judgements about it.
How are AI models Trained?
It's important to understand the fundamental process first before you start dealing with the human element. There are various methods of training AI systems, but a basic one could involve the following steps:
Data > Training > Model > Evaluation > Feedback > Improvement
For a big language model, the primary training may require a huge quantity of text. The model learns statistical patterns that enable it to produce language and undertake a range of tasks.
OpenAI explained that GPT-4's initial training involved it learning to predict the next word from a vast body of information. Further post-training techniques can then be used to steer the model towards desired behaviour.
There are several stages where human involvement can be added.
People may prepare or label data prior to training. They can generate samples for the model to learn on. They can assess its answers following training and evaluate its responses after training. They will rank and order diverse answers to identify errors.
That's why the phrase AI training can encompass a wide variety of activities.
AI Training Is An Umbrella Term
It's essential to understand that AI training is not about a single job role.
Two companies can advertise "AI training" roles that involve completely different tasks.
Depending on the project, you might see titles like:
AI Trainer
AI Training Specialist
AI Evaluator
AI Response Evaluator
LLM Evaluator
AI Data Annotator
Data Labeler
Data Quality Analyst
AI Tutor
LLM Trainer
Search Quality Rater
Prompt Writer
AI Data Specialist
Subject Matter Expert
Domain Expert
AI Safety Evaluator
Certain AI training jobs concentrate on text. Others include pictures, videos, sounds, search results, coding, mathematics, medicine, law, finance or other specific fields.
That is why searching for a job with the broad term as "AI trainer" can lead to a misleading understanding of what kind of work is required.
How Do Humans Help Train AI Models?
There are several ways humans contribute to this process.
1. Data Annotation and Labelling
To learn to identify or label information, AI systems frequently require labelled examples. Humans could, for instance, describe an image as: Car, bus, bicycle, pedestrian.
Text are classified for language systems by sentiment, intent, topic or safety classification. Annotators can label objects in images and video in computer vision. They can translate speech or identify particular sounds in speech projects.
IBM defines data annotation as a human-led process in which humans identify raw data such as images, text files, or videos that can be used for machine learning models.
The exact task will vary with the purpose for which the model is being built.
2. Writing Examples for AI Models
Trainers sometimes need more than just a label. A human may be requested to write an example of a high-quality response.
For insurance, AI system could be given the following prompt:
"Explain inflation to a 12-year old."
A human trainer could write a response that is: Accurate, easy to understand, suitable for audience specified (12 year old) and directly related to the question.
This example can then be used to train the system to show what the system should respond to. It is notable that human-written demonstrations play a significant role in the training of these models.
3. Evaluating AI Responses
Another common form of Human involvement is AI output evaluation.
Here, the trainer is presented with a prompt and a response from the AI and then they are asked to evaluate the response based on some criteria.
For example:
Prompt:
"How to write a professional email / letter to decline a meeting."
The evaluator may ask if the response is: relevant, professional, clear, grammatically correct, complete and appropriate in tone. The evaluator can then score the response or choose one of the categories that describes the quality of the response.
This kind of human evaluation can occur after the model has been trained, as evaluation lets the researchers know how well the model is working and where it needs adjustments. Human evaluations is used alongside with the automated evaluation methods.
4. Ranking AI Outputs
In cases where an AI system generates multiple answers. A human trainer could be asked:
Which is the best answer, A, B or C?
They may rank the responses based on factors like correctness, usefulness, relevance, and safety. This is especially crucial in reinforcement learning from human feedback (RLHF) approaches.
Human labelers were used to compare model outputs in OpenAI's work on InstructGPT. The preferences were then fed into a reward model to learn from and inform the next phase of model optimisation. So, the human does not necessarily "teach the AI one answer at a time" rather, their judgments can give an indication to the system on what behaviors to encourage.
5. Correcting AI Errors
AI systems make mistakes. A human trainer will be tasked with identifying and fixing those errors.
Suppose the answer below is generated by artificial intelligence and it says:
"Sydney is the capital of Australia."
A knowledgeable evaluator would be able to identify the error made and give the correct information: Canberra.
This is the same for more ambiguous and complex activities. A reviewer may rectify:
Factual inaccuracies
Incorrect calculations
Poor grammar
Bad translations
Faulty code
Incorrect reasoning
Misinterpretation of instructions
Inappropriate responses
The corrections may be helpful for enhancing model performance in the future.
6. Domain-Expert Evaluation
Not every task can be evaluated properly by a generalist. Imagine an AI system which answers medical questions.
While a general reviewer might be able to determine if an answer is well written, a qualified medical professional may be better equipped to determine correctness of the clinical information. It applies to every other domain niches like law, finance, engineering, language etc.
That's where subject-matter experts (SMEs) fit into the AI training and evaluation industry.
For instance, ExpertWoka's latest AI job listings feature positions in accounting, economics, law, language, and STEM and other specialized assessment and training. An AI response might be reviewed by a domain expert who can provide additional context, write comprehensive examples, fact-check data, or spot errors in the information that might not be obvious to a less-expert user.
7. Prompt and Response Evaluation.
There is a component of AI training that includes the evaluation of a model's reaction to carefully designed prompts. A contributor could develop or review prompts designed to test particular capabilities.
For example:
Prompt says: Discuss the difference between correlation and causation using a business example.
The evaluator will determine if the AI:
Answered the actual question
Explained both concepts correctly
Cited an example that was suitable
Avoided misleading claims
Followed the requested format
This type of work is used to find weaknesses in the model behaviour.
8. Evaluation of safety and preference
It is possible to ensure that AI outputs comply with certain safety or behavioral standards through human review.
Some evaluators will determine if a response:
- Gives unsafe instructions
- Reveals private information
- Uses inappropriate language
- Complies with the user's proper request
- Withstands a request suitably
- Creates discriminatory, lewd, or harmful materials
The preferences of humans may be included in the process of developing a model, but human preference is not an accurate reflection of all human preferences. This means the quality and consistency of the people that are giving feedback is important.
What is Reinforcement Learning From Human Feedback?
If you have come across the term RLHF, particularly when reading about large language models.
RLHF stands for: Reinforcement Learning from Human Feedback. It's easier to grasp with an example.
Imagine that an AI is presented with a question and generates three answers.
Humans compare those responses, and point out which responses are better based on the project's criteria. Such preferences can be leveraged to create a system to learn which type of outputs humans are likely to prefer. This feedback can then be used to optimize the model.
This approach has been used to improve language model abilities to execute instructions and generate helpful AI responses.
It's important to note that RLHF is one technique not a new name for all AI training.
Does A Human Train The AI Directly?
Usually not in the manner people think. No one sitting behind a computer is altering the internal parameters of the model after each answer. Rather, these human-generated data and feedback may be part of a bigger technical pipeline.
A simplified example is:
Human generates/labels data > collect data > engineers/researchers put the data into a training/evaluation pipeline > model is updated > humans investigate and evaluate results.
The process varies for different organisations and projects. In some systems, human feedback is used to train another model, such as a reward model, which is then used to optimize the primary AI system.
Why Do Human Experts Still Matter as AI Gets Better?
AI models can handle vast amounts of information, but determining if the given answer is actually practical and useful may need context and judgment. That's where Humans comes in. Humans can pick inconsistencies and errors in AI output.
An important caveat is that feedback from humans can be inconsistent or biased. This is because OpenAI's research has confirmed that the behaviour acquired based on human feedback is based on the judgements of the persons giving them feedback. Therefore, good AI training requires clear guidelines, quality control and reviewers that are skilled and qualified.
In an increasingly AI-driven world, human judgment is still a critical component in developing and assessing AI models.
If you are interested in learning about how AI training works, but not what the roles are, then it's time you found out what they look like in practice. Explore carefully curated and verified AI opportunities on ExpertWoka. You can discover the opportunities and see if there's an available position that suits your profile and learn what employers are seeking. It's not essential to be a specialist in AI engineering to grasp the potential role in the AI ecosystem. First, understand the work, and then find opportunities that match what you already know how to do.