Qualia Optimization
This technical report introduced qualia optimization: the problem of considering a candidate measure of an agent's experiential valence alongside task performance. It also proposed representation robustness as a desirable property of candidate valence measures. Representation robustness requires that changing only the representation of a computation does not change the experience attributed to the system. For example, two computers can implement the same abstract computation while using most-significant-bit-first and least-significant-bit-first conventions. Under a computational theory of mind, they should not be assigned different experiences merely because they use different encodings internally.
The report shows that natural measures based on total lifetime reward and cumulative temporal-difference error do not have this property. The same hardware and software can be described using different numerical encodings that assign different values to rewards or TD errors without changing the underlying physical process. These measures can therefore assign different valence to the same physical system solely because its internal quantities are interpreted differently. That conflicts with the physicalist and computationalist assumptions motivating the analysis.
As a possible alternative, the report examines whether the reinforcement of behavior, rather than the magnitude of an internal signal, might be associated with positive valence, and whether inhibition might be associated with negative valence. Likelihood ratios measuring how learning changes the probabilities of past actions provide representation-robust candidates in certain policy-based settings. The report also identifies unresolved issues, including where to draw the agent-environment boundary and how to define reinforcement when an agent does not separate neatly into policy parameters, contextual memory, and distinct learning and acting phases.