Voorbij de Geneesmiddelenontdekking: De Nanotechnologie-Moleculaire Optimalisatie (NMO)-benchmark
Beyond Drug Discovery: The Nanotechnology Molecular Optimization (NMO) Benchmark
June 29, 2026
Auteurs: Matthias Blaschke, Daniel Kienzle, Zsuzsanna Koczor-Benda, Julian Lorenz, Rainer Lienhart, Fabian Pauly
cs.AI
Samenvatting
Generatief moleculair ontwerp wordt gevormd door eenvoudige proxy-benchmarks voor medicijnachtige eigenschappen en modellen die zijn getraind op grote farmaceutische datasets. Deze combinatie levert sterke benchmarkmetrieken op, maar beperkt de overdraagbaarheid naar domeinen die structureel afwijken van medicijnontdekking. Om deze beperking te overwinnen en ontdekking naar echte, wetenschappelijk onderbouwde doelen te sturen, introduceren we de Nanotechnology Molecular Optimization (NMO) Benchmark, die machine learning (ML) en kwantummaterialenwetenschap overbrugt. NMO fungeert tegelijkertijd als een strenge testomgeving voor de ML-gemeenschap en een ontdekkingsmotor voor nanotechnologieonderzoek. De suite vervangt proxy-orakels door kwantumsimulaties en introduceert strikte protocollen die wetenschappelijk nut boven leaderboard-gerichte overfitting prioriteren. De op fysica gebaseerde NMO-taken leggen harde structurele beperkingen en ruige fitness-landschappen op, wat fundamenteel nieuwe eisen stelt aan generatieve modellen. Opvallend is dat geavanceerde moleculaire optimalisatiemethoden veel eenvoudigere benaderingen overtreffen op de NMO-taken. We ontwikkelen een nieuwe basismethode die de kritieke componenten identificeert om de NMO-taken op te lossen, waaronder een nieuwe representatie voor het modelleren van structurele beperkingen en een domein-agnostische pretrainingstrategie om farmaceutische dataset-bias te elimineren. Onze resultaten overtreffen de modernste fysische eigenschappen en onthullen voorheen onbekende structurele motieven, wat nieuwe inzichten biedt voor de nanotechnologiegemeenschap en aantoont dat ML echte wetenschappelijke ontdekkingen kan stimuleren.
English
Generative molecular design is shaped by simple proxy benchmarks for drug-like properties and models pretrained on large pharmaceutical datasets. This combination yields strong benchmark metrics but limits transferability to domains structurally distinct from drug discovery. To overcome this limitation and drive discovery toward real, scientifically grounded targets, we introduce the Nanotechnology Molecular Optimization (NMO) Benchmark, which bridges machine learning (ML) and quantum materials science. NMO acts simultaneously as a rigorous testbed for the ML community and a discovery engine for nanotechnology research. The suite replaces proxy oracles with quantum simulations and introduces strict protocols that prioritize scientific utility over leaderboard-oriented overfitting. The physics-based NMO tasks impose hard structural constraints and rugged fitness landscapes, posing fundamentally new requirements on generative models. Notably, advanced molecular optimization methods underperform much simpler approaches on the NMO tasks. We develop a new baseline method identifying the critical components to solve the NMO tasks, including a novel representation for modeling structural constraints and a domain-agnostic pretraining strategy to eliminate pharmaceutical dataset bias. Our results surpass state-of-the-art physical properties and reveal previously unknown structural motifs, offering new insights for the nanotechnology community and demonstrating that ML can drive genuine scientific discovery.