new

Get trending papers in your email inbox!

Subscribe

Daily Papers

byAK and the research community

Aug 24

Exploring QSAR Models for Activity-Cliff Prediction

Pairs of similar compounds that only differ by a small structural modification but exhibit a large difference in their binding affinity for a given target are known as activity cliffs (ACs). It has been hypothesised that quantitative structure-activity relationship (QSAR) models struggle to predict ACs and that ACs thus form a major source of prediction error. However, a study to explore the AC-prediction power of modern QSAR methods and its relationship to general QSAR-prediction performance is lacking. We systematically construct nine distinct QSAR models by combining three molecular representation methods (extended-connectivity fingerprints, physicochemical-descriptor vectors and graph isomorphism networks) with three regression techniques (random forests, k-nearest neighbours and multilayer perceptrons); we then use each resulting model to classify pairs of similar compounds as ACs or non-ACs and to predict the activities of individual molecules in three case studies: dopamine receptor D2, factor Xa, and SARS-CoV-2 main protease. We observe low AC-sensitivity amongst the tested models when the activities of both compounds are unknown, but a substantial increase in AC-sensitivity when the actual activity of one of the compounds is given. Graph isomorphism features are found to be competitive with or superior to classical molecular representations for AC-classification and can thus be employed as baseline AC-prediction models or simple compound-optimisation tools. For general QSAR-prediction, however, extended-connectivity fingerprints still consistently deliver the best performance. Our results provide strong support for the hypothesis that indeed QSAR methods frequently fail to predict ACs. We propose twin-network training for deep learning models as a potential future pathway to increase AC-sensitivity and thus overall QSAR performance.

  • 4 authors
·
Jan 31, 2023

Monroe: A Molecular Foundation Model for In-Context Probabilistic Inference

Bioassay activity prediction is often data-limited because drug-discovery datasets rely on time-consuming and expensive wet-lab experiments for data generation and evaluation. This challenge has inspired recent research into molecular foundation models (MFMs), which aim to encode general-purpose chemical knowledge into molecular representations that generalize well in data-constrained scenarios. This paper presents Monroe, a new MFM with several innovations over the existing state of the art: increased scale allowing pre-training on over 81 million molecules from the PM6 quantum chemistry dataset; improved graph representation of stereochemistry; improved training losses including conformer denoising and embedding decorrelation; improved multi-task learning; and the use of a prior-data-fitted model (TabPFN) for downstream in-context prediction. Our evaluations use a principled pairwise comparison framework that measures statistically significant performance differences. Across established Polaris benchmarks, Monroe matches or exceeds existing MFMs, while on activity cliff benchmarks, designed to assess utility for molecular discovery, it achieves significant improvements over prior methods. Finally, ablation and transfer experiments show that PFN-based downstream predictors also substantially improve two leading existing models, MiniMol and CheMeleon, yielding new state-of-the-art variants we call MiniMol_PFN and CheMeleon_PFN, suggesting that our downstream adaptation strategy generalizes beyond Monroe. Source code is at github.com/blazejba/monroe.

  • 2 authors
·
Aug 19

LeMMINGs VII: 5 GHz, 50 mas e-MERLIN observations of a statistically complete sample of nearby AGN

We present 5 GHz e-MERLIN radio images at 50 mas resolution of the nuclear regions of the Legacy e-MERLIN Multi-band Imaging of Nearby Galaxies survey (LeMMINGs), the deepest statistically complete radio-band survey of the local Universe (<120 Mpc), consisting of 280 galaxies spanning all morphological and nuclear types. We detect nuclear radio emission above a median 5 sigma threshold of 0.33 mJy beam^-1 in 68 of 280 sources (24 percent), with core luminosities in the range 10^35 to 10^41.9 erg s^-1. The radio emission is attributed to active galactic nuclei, circumnuclear star formation, or, in the case of NGC 3690, a tidal disruption event. The brightest radio nuclei, with brightness temperatures >=10^6 K, reside in optically active galaxies such as LINERs and Seyferts. The detection rate for inactive systems (H II and absorption-line galaxies), which may host low-luminosity active galactic nuclei, is 8 percent. Most detections (78 percent) are compact (<10 pc), while the remaining 22 percent show extended jet-like features up to 380 pc. Compared to the 1.5 GHz LeMMINGs data, the 5 GHz observations provide superior resolution and spatial filtering, resolving out large-scale structures and isolating genuine nuclear emission. Our results suggest that low-luminosity active galactic nuclei are the primary manifestation of black hole activity in the local Universe in the form of compact jets and cores, with a preference for early-type hosts. The two LeMMINGs campaigns indicate that up to 30 percent of the local galaxy population hosts a radio-active nucleus, highlighting the necessity of high-resolution, high-sensitivity imaging for uncovering nuclear emission at the lowest luminosities.

  • 39 authors
·
Mar 8

Scaf-GRPO: Scaffolded Group Relative Policy Optimization for Enhancing LLM Reasoning

Reinforcement learning from verifiable rewards has emerged as a powerful technique for enhancing the complex reasoning abilities of Large Language Models (LLMs). However, these methods are fundamentally constrained by the ''learning cliff'' phenomenon: when faced with problems far beyond their current capabilities, models consistently fail, yielding a persistent zero-reward signal. In policy optimization algorithms like GRPO, this collapses the advantage calculation to zero, rendering these difficult problems invisible to the learning gradient and stalling progress. To overcome this, we introduce Scaf-GRPO (Scaffolded Group Relative Policy Optimization), a progressive training framework that strategically provides minimal guidance only when a model's independent learning has plateaued. The framework first diagnoses learning stagnation and then intervenes by injecting tiered in-prompt hints, ranging from abstract concepts to concrete steps, enabling the model to construct a valid solution by itself. Extensive experiments on challenging mathematics benchmarks demonstrate Scaf-GRPO's effectiveness, boosting the pass@1 score of the Qwen2.5-Math-7B model on the AIME24 benchmark by a relative 44.3% over a vanilla GRPO baseline. This result demonstrates our framework provides a robust and effective methodology for unlocking a model's ability to solve problems previously beyond its reach, a critical step towards extending the frontier of autonomous reasoning in LLM.

  • 7 authors
·
Oct 22, 2025