
under researched gene find
finding under researched gene
Prompt
I want you to identify one highly understudied human gene/protein that could be a strong candidate for a literature review and potentially a future computational biology research project. Primary requirement Find a human gene that has fewer than 20 peer-reviewed research papers specifically focused on that gene/protein's biological function. Do NOT simply count papers where the gene appears somewhere in a large GWAS, RNA-seq, proteomics, cancer panel, or multi-gene study. I want to know the approximate number of papers where the gene is a meaningful subject of investigation. Preferred characteristics Prioritize genes that: 1. Have <20 genuinely relevant research papers. 2. Are poorly characterized or emerging genes/proteins. 3. Have at least some experimentally supported biological function, so there is enough material for a meaningful review. 4. Have a clearly identifiable research gap. 5. Have publicly available data that would allow computational investigation. 6. Preferably have: - structural information/PDB or AlphaFold structure, - proteinâprotein interaction data, - expression data, - RNA-seq/single-cell RNA-seq data, - disease-associated variants, - GWAS/eQTL information, - cancer or cardiovascular datasets, - or other accessible omics datasets. Coronary artery disease preference If possible, prioritize genes associated with coronary artery disease (CAD), atherosclerosis, endothelial biology, vascular smooth-muscle cells, macrophages, lipid metabolism, vascular inflammation, mechanotransduction, or cardiovascular genetics. However, CAD is NOT mandatory. If you cannot find a sufficiently understudied CAD gene, expand to other disease-relevant genes. Important: avoid weak candidates Do NOT recommend a gene merely because it has a low number of publications. Reject candidates that are: - essentially unstudied with no meaningful biological evidence, - pseudogenes unless there is a compelling reason, - poorly annotated genes with no feasible research direction, - genes where almost all evidence consists only of computational predictions, - genes with too little information to support a review, - genes that have actually been extensively studied under an old name/synonym, - or genes whose publication count only appears low because databases use different gene symbols. Also investigate gene synonyms, previous names, aliases, and historical nomenclature before determining the publication count. Search strategy Search broadly across: - PubMed - Europe PMC - Google Scholar if available - NCBI Gene - UniProt - Ensembl - GWAS Catalog - Open Targets - GTEx - Human Protein Atlas - STRING - BioGRID - IntAct - RCSB PDB - AlphaFold DB - ClinVar - gnomAD - relevant cardiovascular databases - recent review articles Use current information available in 2026. Do not rely on a single database. For each candidate Give me: 1. Gene symbol and full name 2. Alternative names/synonyms 3. Chromosomal location 4. Protein length 5. Known/putative function 6. When it was discovered/characterized 7. Approximate number of relevant publications 8. Number of papers specifically focused on the gene, if you can determine it 9. The 5â10 most important papers 10. Whether any paper has experimentally characterized its function 11. Known protein domains 12. Known protein structure/PDB/AlphaFold availability 13. Known interacting proteins 14. Known expression pattern 15. Relevant diseases 16. Whether it has a CAD/atherosclerosis association 17. Known GWAS variants/eQTLs 18. Available transcriptomic/single-cell datasets 19. Known disease-associated variants 20. What is currently unknown about the gene 21. 5â10 specific research questions that remain unanswered 22. 3â5 computational biology projects I could realistically perform 23. Whether those projects could potentially produce a publication 24. What experimental validation would ideally be performed later 25. Why this gene is more interesting than other understudied genes Most important part: novelty assessment For the top candidates, explicitly answer: ÂŤâWhy is this gene worth studying in 2026?âÂť Identify what has changed recently and whether the gene has suddenly become relevant because of: - a new structural paper, - a new GWAS association, - a new disease connection, - a new single-cell dataset, - a new protein interaction, - a newly discovered pathway, - or a newly characterized molecular function. I am particularly interested in genes where a small number of existing papers collectively suggest an important biological role, but the mechanism is still poorly understood. Rank the candidates Find at least 10 candidates initially, then rank the best 5 using: Criterion| Weight Genuine publication scarcity| 25% Biological importance| 20% Strength of existing evidence| 15% Research gap/novelty| 15% CAD/vascular relevance| 10% Availability of computational data| 10% Potential for future publication| 5% Give each candidate a score out of 100. Final recommendation At the end, select ONE gene as your #1 recommendation. For that gene, provide: - a concise literature timeline, - the key papers, - what is established, - what is uncertain, - the biggest research gaps, - the most promising computational project, - and a proposed review-paper title. Critical requirement Be extremely careful with the â<20 papersâ criterion. For the final recommended gene, show me exactly how you arrived at the publication estimate, including: - search terms used, - synonyms searched, - databases checked, - approximate number of results, - which papers were excluded and why, - and the final number of genuinely relevant papers. I would rather receive fewer candidates with verified publication scarcity than a long list of genes whose publication counts are inaccurate. Do not manufacture publication counts, papers, DOI numbers, disease associations, or experimental findings. Clearly distinguish between experimentally demonstrated, reported association, computational prediction, and your inference. Finally, tell me whether the candidate is suitable for: (A) a literature review, (B) an undergraduate computational biology project, (C) a potential research paper, and (D) future wet-lab validation.