
Developing a new drug is a long and expensive process, which often takes 10-15 years and costs more than US$ 2 billion (Berdigaliyev & Aljofan, 2020). It contains three main steps: Understanding the disease and identifying a treatment target, designing an appropriate drug and testing it in clinical trials. Each step is composed, time-consuming and resource intensive, delays the arrival of new treatments.
Artificial intelligence (AI), especially large language model (LLM), provides promising methods for speeding up this process. LLMs can understand scientific texts, support disease research, help with drug design and clinical testing can help handle data. Recent models have been used for tasks such as identifying the goals of the drug, automating chemical experiments and analysis of clinical knowledge. With further progress, LLM drugs can play an important role in all stages of development, which can help improve efficiency and results.
2. The Main Type of Language Model in Drug Development
In the discovery of the drug, scientific languages as a smile (used to represent molecules) and fixed (used for protein and genetic sequences) are crucial to code biological information. To interpret these formats effectively, two main types of language models are used: special models and general objective models.
2.1 Specialized Large Language Model
Specialized LLM is trained on scientific data and is designed to handle specific languages used in biology and chemistry. These models can extract meaningful patterns from raw researchers.
In disease research, they help by analyzing gene expression data and DNA sequences, and revealing an insight such as genetic markers, regulatory elements and genetical networks. Some protein-based models can predict the structure or function of the protein by analyzing amino acid sequences and improving understanding of the disease system.
In the discovery of the drug, these models help predict chemical reactions, plan synthesis and suggest new molecules with desired properties. They also support the first safety screening by predicting absorption, metabolism, and toxicity (ADMET).
Usually, these models work to take defined inputs as a protein sequence and a molecule SMILES code and return a prediction, such as binding power or the possibility of interaction.
2.2 General purpose Large language model
General Purpose LLM is trained on a wide range of lessons, including scientific literature, so that they can understand complex biomedical materials and apply it to different fields.
They can review the huge versions of the published data, summarize the findings and create a relationship between genes and diseases, and support the target identity. These models also help explain technical words in single languages, making it easy to understand and communicate scientific material.
In experimental chemistry, the general LLM has shown the ability to automate functions such as prediction and synthesis scheme. Although they cannot match the special model for each task, ongoing research seeks how to generate or modify proteins and molecules design using extensive scientific knowledge.
In clinical research, general LLM patients help undergo records, design studies and mating patients with appropriate tests. Their ability to analyze documents and generate summaries supports planning and documentation in clinical studies.
Users often interact with a general model who uses natural language, questions and receives reactions based on existing knowledge in literature.
LLM in Drug Discovery and Development
Large language models (LLM) become valuable tools in findings and development of medication. Their ability to treat and interpret gigantic and complex biological data sets allows researchers to extract meaningful insights that can accelerate and increase different stages of the drug development process.
In understanding the disease system, LLM genes can analyze scientific literature, clinical records and large versions of omics data to identify associations between genes, proteins and disease phenotypes. This helps researchers generate hypothesis about how diseases occur at molecular level and progress.
In genomics and transcriptomics, LLM genes help explain highly expanding sequencing data, revealing gene expression and patterns in regulatory elements. This insight can be important to identify biomarkers and medical goals. LLMs can also help prioritize genes based on their relevance to a particular disease, which may enable more concentrated and effective research.

When it comes to protein target analysis, LLM can treat structural and functional protein data to predict medication and interaction sites. This drug supports the identity of viable goals for development. In combination with molecular simulation and other AI units, LLMs can also help in model protein-ligand interactions that help the design of potential medical science.
Overall, the integration of LLM in the discovery of the drug in the early stage provides a more computer-driven and future staging, which helps researchers to reduce test-and-tough processes and accelerate the development of new means.
Maturity Assessment of Large Language Models in Drug Development
This section explains how large language models (LLMS) develop in various tasks in pharmaceutical development pipelines. The evaluation is divided into three main stages: the disease system, the discovery of the drug and understanding of the clinical studies. Each step consists of specific downstream functions where LLMs are quickly used.

Four-level scale is used to evaluate the maturity of LLM use in each task:
• Level 0 (no capacity): LLM has not yet shown meaningful results for this task.
• Level 1 (Basic Performance): LLM has been tested in academic surroundings, but lacks evidence from real-world applications.
• Level 2 (new use): Early stages exist, often in the industry, although extensive use is still limited.
• Level 3 (advanced and valid): LLMs are usually used in real-world scenarios and are supported by practical verification.
Below is a breakdown of maturity level in 14 downstream functions.
1. Understand the Disease System
• Genomic Analysis: Specialized LLM designed for genomic functions is well installed. Many studies have shown their value in predicting gene function and discovering mutations, keeping this task at the most advanced level.
• Transcriptomics Analysis: Models used in this field are still developing. Most are in the practical phase and not yet used to display continuous results in settings.
• Protein Target Analysis: This task has seen rapid progress, especially after the release of Alphafold2. The unit has a fairly advanced protein structure prediction and is widely available, useful for supportive functions such as drug design and vaccine research. As a result, LLM in this region is considered very mature.
2. Drug Discovery
• Goal Identification: LLM is used to match biological goals with diseases. While promising models are available, the use of use in the industry is still an early stage and is not yet wider.
• Hit Discovery and Lead Optimization: These features include the choice and processing of promising compounds. LLM has shown some success, especially when combined with other AI units. However, these methods are still developing and practical use and require confirmation.
• ADMET Prediction: It is necessary to predict the development of drug to predict absorption, distribution, metabolism, emissions and toxicity. Many LLM models are available for this purpose, and their use is increasing. They come closer to the more stable and practical development phase.
• De Novo Drug Design: Generic models that use LLM are discovered to create new molecules. While the field is of strong interest, technology is still in development, and the results of the real-world medicinal pipelines are limited.
• Synthetic Pathway Prediction: This involves prediction how to create a compound when using chemical reactions. Some LLM shows promises here, but the region is still in the early stages of the adoption.
3. Clinical Trials:
• Patient Recruitment: LLM can help find suitable test participants by analyzing health records and other data. Although capacity is clear, the use of these models is still developing in actual tests.
• Trial Design: LLM is used to help design more effective tests. The current models are still tested, and practical applications are limited.
• Protocol Generation: LLMs can help prepare the clinical protocol. Although it is useful to reduce time and effort, human monitoring is necessary. Adoption increases, but is not widespread yet.
• Data Analysis: Analysis of test data using LLMs draws attention. However, the model must meet strict standards for accuracy and compliance, which has slowed their widespread use.
• Adverse Event Prediction: Predicting possible security problems before it is a complex task. LLM is trained to support this process, but the technology is still in the early stages.
Future Direction
As the Large Language Models (LLMs) develops, their role in the development of the drug is expected to increase. Future efforts will focus on improvement in model accuracy, reliability and openness. This involves reducing errors in predictions and making the model output easier, which is especially important in safety-creating areas such as health services.
Another large development area includes the construction of domain-specific LLM. These models are trained on special biomedical data and correspond to the discovery of the drug or the special stages of clinical research. Such targeted models are more likely to produce relevant and accurate results than normal-cleaned LLM.

Integration with other technologies, such as screening high throughput, laboratory automation and computer systems in the real world, are also expected. These combinations can help improve the efficiency and success rate of drug growth by providing rapid insight and reducing manual work.
However, broad adoption will require more evidence from real applications. There is also a need to address problems such as privacy, moral use and regulatory acceptance. Cooperative efforts between researchers, industries and regulators will be required to go from experimental use to regular practice.
In the long term, LLMS can become a standard part of pharmaceutical research and development, which supports several tasks from goal search to test monitoring. Their success will not only depend on technological progress, but also on responsible distribution and continuous evaluation.