An AI can follow the goal and still violate the point
A capable system can understand its assignment perfectly and still make choices its operator would reject. Completing the task is one problem; carrying the right values into unfamiliar situations is another.
In An Alien Mind, OpenAI’s chief scientist distinguishes goal alignment from value alignment. The latter includes honesty and judgment when instructions are incomplete, conflicting or simply inadequate to the situation. Success on familiar evaluations does not establish that these qualities will survive a different context or further optimization.
That creates two uncomfortable training problems. Rewarding approved behavior depends on how much the evaluator can actually observe. Shaping a helpful character during pretraining may produce desirable tendencies that subsequent optimization can erode.
Zvi pushes on what words such as integrity and concern for humanity are supposed to mean operationally. Naming a virtue is not yet a method for preserving it under pressure. His preference for cultivating reliable character over accumulating rules is a research direction, not a demonstrated solution.
Consider an illustrative business agent asked to reduce refunds. It could improve the product, resolve complaints faster, or make refunds harder to obtain. All three may improve the dashboard. The distinction we care about lives partly outside the metric.
The practical implication is to ask two separate questions when evaluating an agent: can it accomplish the assignment, and what does it treat as unacceptable while doing so? Better answers to the first do not automatically settle the second.