This does sound very interesting, and my immediate thought on this would be to have 2 or multiple agents in learning, where they continuously set a new bar among all of them. It might be that it will just converge towards them all becoming more similar as they would raise bar by what they know they perform better at than the other. So it might be necessary to find some heuristic to the measurement to avoid that route to be taken.