Exploring chemical space for lead identification by propagating on chemical similarity network

Yi, Jungseob; Lee, Sangseon; Lim, Sangsoo; Cho, Changyun; Piao, Yinhua; Yeo, Marie; Kim, Dongkyu; Kim, Sun; Lee, Sunho

Detailed Information

Cited 4 time in webofscience

Cited 5 time in scopus

Metadata Downloads

Exploring chemical space for lead identification by propagating on chemical similarity networkopen access

Authors: Yi, Jungseob; Lee, Sangseon; Lim, Sangsoo; Cho, Changyun; Piao, Yinhua; Yeo, Marie; Kim, Dongkyu; Kim, Sun; Lee, Sunho

Issue Date: Jan-2023

Publisher: Elsevier B.V.

Keywords: Chemical network construction; Data mining; Lead identification; Network propagation

Citation: Computational and Structural Biotechnology Journal, v.21, pp 4187 - 4195

Pages: 9

Indexed: SCIE
SCOPUS

Journal Title: Computational and Structural Biotechnology Journal

Volume: 21

Start Page: 4187

End Page: 4195

URI: https://scholarworks.dongguk.edu/handle/sw.dongguk/21057

DOI: 10.1016/j.csbj.2023.08.016

ISSN: 2001-0370
2001-0370

Abstract: Motivation: Lead identification is a fundamental step to prioritize candidate compounds for downstream drug discovery process. Machine learning (ML) and deep learning (DL) approaches are widely used to identify lead compounds using both chemical property and experimental information. However, ML or DL methods rarely consider compound similarity information directly since ML and DL models use abstract representation of molecules for model construction. Alternatively, data mining approaches are also used to explore chemical space with drug candidates by screening undesirable compounds. A major challenge for data mining approaches is to develop efficient data mining methods that search large chemical space for desirable lead compounds with low false positive rate. Results: In this work, we developed a network propagation (NP) based data mining method for lead identification that performs search on an ensemble of chemical similarity networks. We compiled 14 fingerprint-based similarity networks. Given a target protein of interest, we use a deep learning-based drug target interaction model to narrow down compound candidates and then we use network propagation to prioritize drug candidates that are highly correlated with drug activity score such as IC50. In an extensive experiment with BindingDB, we showed that our approach successfully discovered intentionally unlabeled compounds for given targets. To further demonstrate the prediction power of our approach, we identified 24 candidate leads for CLK1. Two out of five synthesizable candidates were experimentally validated in binding assays. In conclusion, our framework can be very useful for lead identification from very large compound databases such as ZINC. © 2023 The Author(s)

Files in This Item: There are no files associated with this item.

Appears in Collections: College of Advanced Convergence Engineering > Department of Computer Science and Artificial Intelligence > 1. Journal Articles

Show full item record

qrcode

Related Researcher

Researcher Lim, Sang Soo photo

Lim, Sang Soo: College of Advanced Convergence Engineering (Department of Computer Science and Artificial Intelligence)

Read more

Altmetrics

Total Views & Downloads

RSS_1.0 RSS_2.0 ATOM_1.0

30, Pildong-ro 1-gil, Jung-gu, Seoul, 04620, Republic of Korea+82-2-2260-3114

Certain data included herein are derived from the © Web of Science of Clarivate Analytics. All rights reserved.
You may not copy or re-distribute this material in whole or in part without the prior written consent of Clarivate Analytics.

Detailed Information

Related Researcher

Altmetrics

Total Views & Downloads

BROWSE