I am a computer scientist and a computational intelligence researcher in the field of swarm intelligence. The contemporary terms 'AI safety' and 'AI security' do not fall within my vocabulary; rather, I simply consider that the performance of LLMs should be evaluated using classical measures like precision and recall, along with various other statistical measures.
Actually, new measures are being published day by day in academia and the business sector. Meanwhile, the discussion on how humans use such statistics effectively is not satisfactory. Go back to the fundamental statistics that involve human evaluation in their use, and go forward to the next step—that is my belief.
No comments:
Post a Comment
Note: Only a member of this blog may post a comment.