# Best approach for L2G mapping with GWAS variant lists for a specific indication

**URL:** <https://community.opentargets.org/t/best-approach-for-l2g-mapping-with-gwas-variant-lists-for-a-specific-indication/1939>\
**Category:** Data Access\
**Created:** [4 November 2025 12:00 UTC](https://community.opentargets.org/t/best-approach-for-l2g-mapping-with-gwas-variant-lists-for-a-specific-indication/1939 "2025-11-04T12:00:28Z")\
**Posts on this page:** 2\
**Page:** 1

<div class="post-metadata">

**Author:** ![jkozlowska](https://avatars.discourse-cdn.com/v4/letter/j/7cd45c/32.png) [@jkozlowska](https://community.opentargets.org/u/jkozlowska)\
**Post date:** [4 November 2025 12:00 UTC](https://community.opentargets.org/t/best-approach-for-l2g-mapping-with-gwas-variant-lists-for-a-specific-indication/1939/1 "2025-11-04T12:00:28Z")

</div>

Hi there!

I have long lists of GWAS-associated variants and would like to obtain L2G scores to identify likely causal genes. I’d appreciate guidance on the best approach as it’s currently unclear. I’ve prepared my data in parquet/CSV files with these columns:

- `variant_id`: chr:pos:ref:alt (e.g., “4:79957426:C:T”)

- `rsid`: dbSNP identifier (e.g., “rs10000339”)

- `chromosome`, `position`, `reference_allele`, `alternate_allele`: Genomic coordinates (GRCh38)

- `gwas_hit`: Lead variant for the locus

- `p_value`: Association p-value from GWAS

- `indication`: Disease indication (e.g., “uc”)

These are GWAS variants plus variants in LD (not fine-mapped credible sets). My Questions

1. Can I use existing Open Targets L2G scores?

Is there a way to query/download L2G scores for my specific variant list? I see L2G data in the Open Targets Genetics portal, but I’m unclear on:

- Can I batch query with rsIDs or variant IDs?
- What’s the recommended approach: API, GraphQL, FTP download?

1. Do I need to/can I run L2G locally via gentropy?

I’ve looked at the gentropy CLI documentation for `locus_to_gene` step, but I notice it expects:

- Credible sets
- Feature matrices
- Reference datasets

Would L2G work on raw GWAS variants or is finemapping necessary for the pipeline to work?

1. What’s the recommended workflow?

Given my starting point (GWAS variants + p-values only), what would you recommend:

Option A: Query existing Open Targets L2G scores

- Map my variants to OT study loci
- Download pre-computed L2G predictions
- Filter for my indications

Option B: Run fine-mapping first, then L2G

- Use SuSiE/FINEMAP to create credible sets
- Run gentropy L2G locally
- More rigorous but time-intensive

Or is there something else I’m not considering?

Any guidance on the best approach would be greatly appreciated!

---

<div class="post-metadata">

**Author:** ![Szymon\_Szyszkowski](https://dub1.discourse-cdn.com/flex017/user_avatar/community.opentargets.org/szymon_szyszkowski/32/757_2.png) [@Szymon\_Szyszkowski](https://community.opentargets.org/u/Szymon_Szyszkowski)\
**Post date:** [10 November 2025 13:22 UTC](https://community.opentargets.org/t/best-approach-for-l2g-mapping-with-gwas-variant-lists-for-a-specific-indication/1939/2 "2025-11-10T13:22:36Z")

</div>

Welcome to the community @jkozlowska.  
Before I can answer your (many) questions, lets prepare the background and summarise all of the knowledge behind L2G that is currently implemented on the Platform.

## L2G steps breakdown

As you have found, to obtain the _L2G scores_ we use [gentropy steps](https://opentargets.github.io/gentropy/python_api/steps/_steps/). The whole process of obtaining the scores is divided into 3 parts:

1. Building L2G Feature Matrix
2. Training model
3. Predicting scores

I had made a short schema of the process below

 ![image](https://europe1.discourse-cdn.com/flex017/uploads/opentargets/original/1X/efeae8bd1a2f7b27b376f9d5f78f954e56bce5e0.png)

All dataset paths are written in red and are relative to the `https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/latest` release.  
This architecture is written in our [Unified Pipeline configuration](

## Answers

Given the knowledge above, I can now answer your questions directly

> 1. Can I use existing Open Targets L2G scores?

Yes, you can use the GraphQL api to **retrieve scores for specific variant**.

```auto
{ 
  variant(variantId: "4_79957426_C_T") {
    id
    GWASCredibleSets: credibleSets(
      studyTypes: [gwas]
    ) {
      count
      rows {
        studyLocusId
        studyId
        l2GPredictions {
          rows {
            score
            target {
              id
            }
          }
        }
      }
    }
  }
}

```

This query should return the variant you were looking for and it’s **L2G scores across all credible sets**.

If you want to **get the scores for loci associated with the disease of interest** (assuming uc is [_Ulcerative Colitis_](https://platform.opentargets.org/disease/EFO_0000729) ), you need to search for studies linked to the disease

```auto
{
  studies(diseaseIds: "EFO_0000729") {
    rows {
      id
    }
  }
}

```

and finally filter results from the first query (all credible sets) by the studyIds from the second query. (not shown).

**Note that the L2G score is not describing variant, rather full credible set.** To have a best proxy of variant score estimation you would need to make sure that your variant has a high (\>0.9) _posterior inclusion probability_ within the searched credible set.

> Is there a way to query/download L2G scores for my specific variant list? I see L2G data in the Open Targets Genetics portal, but I’m unclear on:
> 
> - Can I batch query with rsIDs or variant IDs?
> - What’s the recommended approach: API, GraphQL, FTP download?

The preferred way to query in batch is to download the datasets

- [predictions](https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/latest/output/l2g_prediction/) to find `studyLocusId` and L2G scores
- [credible\_set](https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/latest/output/credible_set/) to bring `variantId` that account to `studyLocusId`
- [study](https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/latest/output/study/) to filter out `studyLocusId` by the relevant disease

Alternatively you can use [big query](https://platform-docs.opentargets.org/data-access/google-bigquery) to run SQL directly on the datasets.

> 1. Do I need to/can I run L2G locally via gentropy?

After querying if you can not find your variants, to obtain L2G scores for them you would eventually need to run the fine-mapping (fortunately with in-sample-LD) and obtain credible sets.

To obtain the predictions for **new credible sets** You need to:

- Transform your credible sets to [StudyLocus format](https://opentargets.github.io/gentropy/python_api/datasets/study_locus/#schema).
- Build [L2G Feature Matrix Step](https://opentargets.github.io/gentropy/python_api/datasets/study_locus/#schema) to obtain credible sets.
- Run the [L2G step](https://opentargets.github.io/gentropy/python_api/steps/l2g/#gentropy.l2g.LocusToGeneStep) with generated feature matrix and credible sets using [pre-trained model](https://ftp.ebi.ac.uk/pub/databases/opentargets/platform/latest/etc/model/locus_to_gene_model/)

### Building Feature Matrix

As you may have seen from the diagram above, to generate the L2G feature matrix, one need to provide:

- colocalisation dataset that contains results from colocalising **your GWAS credible sets** to QTL credible sets
- variant index that contains all variants from **your GWAS credible sets** annotated with VEP
- study index dataset that contains information about the molQTLs studies (the column `geneId` is used to determine molecular feature affected by the credible set linked to study
- target index dataset - this can be used directly from the platform output without any modifications

> Would L2G work on raw GWAS variants or is finemapping necessary for the pipeline to work?

No, the feature matrix step requires the `PIP` to be present for each individual variant in locus as it is used as a feature weight, so **if you have variants, you need to run the fine-mapping before building L2G feature matrix.**

> 1. What’s the recommended workflow?  
> Given my starting point (GWAS variants + p-values only), what would you recommend:  
> Option A: Query existing Open Targets L2G scores
> 
> - Map my variants to OT study loci
> - Download pre-computed L2G predictions
> - Filter for my indications  
> Option B: Run fine-mapping first, then L2G
> - Use SuSiE/FINEMAP to create credible sets
> - Run gentropy L2G locally
> - More rigorous but time-intensive

As mentioned above, I would start from **querying existing predictions**.

If you would not find the variants you are looking for, then **I would suggest running fine-mapping and building the feature matrix, running L2G model with both.**.

By looking at your workflow ideas, I feel that you know most of the stuff already, if there is anything unclear, feel free to raise questions.

With kind regards,  
Szymon Szyszkowski
