Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists — arXiv cs.AI. arXiv:2608.12345v1 Announce Type: new Abstract: Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct classificati…
Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments — arXiv cs.AI. arXiv:2608.12368v1 Announce Type: new Abstract: Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents…
Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese — arXiv cs.AI. arXiv:2608.12373v1 Announce Type: new Abstract: Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can…