AI Ethics: Understanding Human Values in AI

A Large Language Model (LLM) undergoes several training steps:
- Pre-Training (self-supervised, trained on internet data)
- Post Training (Instruction Finetuning & Reinforcement Learning from Human Feedback)
In the pre-training phase, the model learns from vast internet data, good and bad, producing a powerful but raw language model, not ready for commercial use. In post-training, the model is instruction-finetuned (supervised) to behave like a chatbot. While the LLM can now follow prompts, it may still generate unsafe outputs. A second post-training step is required and focuses on alignment, tuning the model to reflect human values and ensure safe behavior.
The alignment problem is the challenge of ensuring AI systems behave in line with human values, ethics, and intentions. This is especially difficult, as human values are subjective and context-dependent. Trained on internet data, LLMs often inherit biases and contradictions, which can lead to harmful or misleading outputs if not properly addressed.
Only when the model is sufficiently aligned with human values, the LLM vendor will release the model for commercial use.
Challenges
AI companies are constantly balancing the need to prioritize safety with the drive to accelerate model commercialization, fuelled by increasing commercial pressures.
- Safety & Quality vs. Speed: In May 2022, before the release of ChatGPT by Open AI, a startup, Anthropic, chose to delay the release of its chatbot (Claude), to continue internal safety testing. This decision was made despite the opportunity to secure a first-mover advantage, highlighting the company's commitment to responsible AI deployment. AI companies continuously juggle prioritizing safety and the drive for rapid model commercialization, driven by mounting commercial pressures. This decision likely cost Anthropic billions of dollars.
- Robust Alignment: Aligning LLMs remains a challenge as jailbreak techniques keep bypassing safety measures. Earlier this year, OpenAI rushes to ban GodMode an app that teaches users 'how to hotwire cars and cook meth at home'. The alignment problem requires methods that address deeper, systemic issues rather than surface-level patches.
- Mechanistic Interpretability: Understanding how AI models reason is vital. Mechanistic interpretability seeks to reveal how models generate outputs, aiding in guiding their behaviour. Progress is ongoing but this remains an open challenge.
- Open-Ended Learning and Self-Improvement: AI Companies are exploring the potential of open-ended learning as a way to achieve better alignment. These Agentic AI systems, which continuously generate and learn from new tasks, are seen as a pathway to developing more general and adaptable AI systems. Agentic AI however, presents new challenges, particularly tied to the "learning by doing" concept.
- Handling Complex Trade-offs: LLMs may face trade-offs between competing human values. For example, providing a response that maximizes accuracy might conflict with the need to be non-offensive or inclusive. Navigating these trade-offs requires sophisticated decision-making.
Reinforcement Learning from Human Feedback (RLHF)
One of the most common methods for aligning LLMs with human values is RLHF and this technology was used by Open AI to humanize the outputs of ChatGPT. Here's how it works:
The first step in AI model training for an LLM like GPT-4 is pre-training. This involves using a self-supervised learning approach on an extensive dataset, such as the internet. For instance, consider a small excerpt from the EU AI Act. In pre-training, approximately 15% of the words or sentences are masked, and the model learns to predicting these masked elements. Since no explicit and expensive human labelling is required, this process is called self-supervised learning. The outcome of the pre-training process is a numerical representation of language also known as a language model. A language model is called 'large' when it has over 1 billion parameters. The P in GPT-4 stands for pre-trained. At this stage, the model's output still feels rough and inconsistent. It is worth noting that the compute costs of pre-training LLMs are/were substantial. Rumour is that OpenAI spend over USD 100 million for pre-training GPT-4.
Instead of training massive models, some smaller models are trained by learning from larger ones. DeepSeek reportedly used this method to train competitive models on fewer GPUs. This technique is called model distillation. Distillation can be seen as ethical because it improves accessibility by making large AI models smaller, faster, and usable on everyday devices. It also reduces energy consumption, which benefits the environment, and typically reuses existing training data, avoiding new privacy risks.
However, concerns arise when distillation is used to replicate proprietary models without permission, raising intellectual property issues. The distilled version might also behave differently, making it harder to ensure safety, fairness, or accountability. Additionally, the process can reduce transparency, making it more difficult to understand how decisions are made.
In the end, whether distillation is ethical depends on how it’s done, who does it, and what the final model is used for.

After pre-training the LLM may be instruction-finetuned on specific tasks. The pre-trained LLM is refined using task-specific data designed as instructions. It uses a supervised learning approach, where models are trained on instruction-response pairs curated by humans. After instruction-finetuning, the LLM's response quality will improve but it will still not feel 100% safe.
Infusing complex human values into the LLM via supervised learning is an impossible task, as we would need to curate millions of examples to serve as training data. Remember that all Human-in-the Loop (HITL) activities increase the cost of training substantially.
The Reward Model
We can train a model that is able to score outputs, just like humans would do. The model would provide a score, number between 0 and 1, associated with an answer to a specific question (prompt).

The issue however is that human feedback scores are subjective, hence noisy. Instead of scoring answers directly, a human is given the simple task to rank two answers given a prompt. Repeated ranking trains a reward model that knows how to score answers like humans.
Fine-tuning with RLHF: The overall training scheme for RLHF is as follows:

Before we start our reinforcement learning process to align the LLM with human values, we must first train the reward model and freeze its weights. Subsequently we start RLHF training and a prompt is given to the 'to be trained' LLM (Where is Brussels). The LLM provides an answer (In Belgium). The answer is fed into the reward model that scores the answer (0.96). The score is given back to the LLM, that will adjusts its weights. After alignment the model is ready for commercial use.
The RLHF process continues after the LLM is commercialized. The thumbs up/down symbol reinforces ChatGPT's answer.

Given the complexity of RLHF, researchers are exploring more efficient methods.
- DPO: Direct Preference Optimization focuses on directly modeling human preferences through a simpler optimization process. DPO offers lower computational costs and has a lower risk of reward hacking.
- Constitutional AI: Constitutional AI, pioneered by Anthropic, uses a predefined set of principles, called a 'constitution" to guide the AI's behaviour and ensure its outputs are safe, ethical, and aligned with societal norms. The methodology however requires carefully crafted guidelines that may not adapt well to more complex human preferences.
- Few-Shot Prompting: Few-Shot Prompting is fast, cheap and leverages the in-prompt/in-context learning feature of LLMs. A carefully engineered prompt provides a few examples at inference time, guiding the output. This performance however depends heavily on prompt quality.
- Adversarial Training: In adversarial training, an additional model (adversary) is used to test the LLM by generating challenging examples designed to expose misaligned behaviour. Adversarial training improves model robustness but it is resource-intensive.
- Human-in-the-Loop (HITL): Human-in-the-Loop methods involve continuous monitoring and feedback from human evaluators. The model is updated periodically based on new human feedback to improve alignment iteratively. HITL evaluation is dynamic and adaptable to new preferences, but costs are high as humans involved to train the reward model
Understanding Safety Challenges in Practice
Enterprises and SMBs (Small Medium Businesses) that use these LLMs on a daily basis, have little impact on the alignment process itself as most models comes aligned. Not all models have been aligned in the same way though. Anthropic for example doesn’t use traditional RLHF. Instead, Anthropic trains Claude using Constitutional AI, a form of RLAIF (Reinforcement Learning from AI Feedback). The key point is that feedback comes from the LLM itself, guided by a predefined constitution, rather than from human reviewers. This alignment method may become a crucial factor when choosing an LLM for specific applications. Ultimately, each company remains accountable for any AI-generated content it produces or distributes. Companies must recognize that LLMs, our AI Agents, do not possess reasoning abilities or consciousness.
The complexity of AI alignment becomes particularly evident when organizations attempt to deploy AI systems at scale. Issues around AI consciousness misconceptions and copyright infringement risks highlight why proper alignment training is critical for enterprise deployments.
Conclusion
The alignment problem remains one of the most critical challenges in LLM model development, requiring a delicate balance between ethics, safety, and commercial viability. While approaches like RLHF, Constitutional AI, and adversarial training show promise, they underscore the complexity of embedding subjective human values into objective algorithms.
As AI systems become more sophisticated and autonomous, the importance of robust alignment methodologies will only increase, particularly for enterprise applications where AI agents interact directly with customers and business processes.
Rainmakers SG helps small and medium businesses design safe and scalable Agentic AI systems that provide immediate ROI!



