Advertisement
Sections
Open AI research shows that human feedback can speed up machine learning tasks
Human feedback can complement existing AI development approaches such as imitation learning and reinforcement learning.

Research by Open AI has demonstrated the potential of human feedback to accelerate machine learning. The researchers have developed an algorithm that can take a guess at what humans want, based on human feedback on which of two behaviours is better. The researchers used the new algorithm to develop a backflip, by letting humans select the better option between two choices. This approach is better than specifying goals or writing a reward function can lead to undesirable behavior from the AI.The algorithm was used for a series of Atari games as well. The machines learnt to anticipate human preferences based on the feedback, instead of the goals in the game. In Seaquest, the agent learned to value oxygen, and worked out how to recover from crashes in the racing title Enduro. The performance of the agents matched the humans, and at times the agents even developed superhuman capabilities. The humans could provide the agents with feedback not aligned with the goal of the environment, so the agents could match the score of another car in Enduro instead of beating the score.
The aim of the research is to improve the safety of AI solutions. This approach of human feedback can complement existing AI development approaches such as imitation learning and reinforcement learning. Using human feedback to train the machines allows for advancing the capabilities of the agent much faster than manually hand crafting the objectives.
The aim of the research is to improve the safety of AI solutions. This approach of human feedback can complement existing AI development approaches such as imitation learning and reinforcement learning. Using human feedback to train the machines allows for advancing the capabilities of the agent much faster than manually hand crafting the objectives.First Published:Jun 15, 2017, 09:56:11 IST
Advertisement
Advertisement

Why AI notetakers are raising serious privacy and security concerns
AI notetakers promise effortless meeting summaries, but experts warn they could expose confidential conversations, corporate secrets and personal voiceprints. As businesses increasingly adopt AI-powered meeting assistants, questions over data storage, privacy, consent and legal risks are becoming impossible to ignore
5 min read
China's low-cost AI models are changing the global AI race. Here's why Silicon Valley is worried
As Chinese firms continue to improve performance while keeping prices low, the AI race is no longer just about building the smartest model—it is increasingly becoming a battle over who can deliver the best value
2 min read
China's Kimi K3 challenges US AI leaders with frontier-level performance at lower cost
Chinese artificial intelligence startup Moonshot AI has unveiled its latest open-weight AI model, Kimi K3, with early results suggesting it could compete with some of the world's most advanced AI systems developed by leading US companies
2 min read
How did Instagram run ads promoting child abuse in India?
India has issued a notice to Meta after an investigation alleged that Instagram displayed paid advertisements promoting child sexual abuse material. MeitY ordered Meta to remove such Instagram ads and explain within seven days how they were approved
8 min read
Why has India halted WhatsApp’s username feature before launch?
India has halted WhatsApp’s planned username feature, citing concerns about cybercrime, impersonation, and law enforcement challenges. As Meta races to address security concerns, here’s why MeitY has paused the rollout, what the feature does, and how it could influence privacy and online safety in India
3 min read
Advertisement
Advertisement
