Model Description

The ASMTransformers model is a machine learning model that calculates embeddings of assembly functions to enable comparison and similarity operations. Instruction set architectures (ISAs) ARM, AMD64, RISC-V and i386 assembly are supported. The aim is to aid in the process of reverse engineering.

Status

Not actively maintained after publishing

Relevant links

Version

2026-08-06

Usage

To use this model, use inference.py in our repository.

Intended use

The model is trained and tested to be used for similarity search of assembly code. It is not trained/tested on any other languages than ARM, AMD64, RISC-V or i386, nor is it tested on other downstream tasks.

Architecture description

The model is a Bert-base architecture with 12 layers, 12 attention heads and a hidden size of 768. It uses a vocabulary of size 6161 derived from a custom assembler architecture-aware tokenizer.

The model is trained in two phases - a masked language modeling (MLM) pretraining phase and a fine-tuning phase for function similarity.

The MLM pre-training phase uses a loss function that includes jump target prediction (see Wang et al.). This pretraining is perfomed using a masking probability of 0.4 that was determined through ablation experiments. This is in line with the findings of Wettig et al., 2023.

We perform the finetuning phase using Batched Semi Hard Triplet Loss on batches of size 1024. The training set for this phase includes triplets that are formed using versions of the same function compiled for different instruction set architectures.

Estimated training time:

pretraining: 38 hours on 4 NVIDIA H200s finetuning: 4 hours on 1 NVIDIA H200

Full architecture:

[
  {
    "idx": 0,
    "name": "0",
    "path": "0_ASMTransformerModule",
    "type": "asmtransformers.models.asmsentencebert.ASMTransformerModule"
  },
  {
    "idx": 1,
    "name": "1",
    "path": "1_Pooling",
    "type": "sentence_transformers.sentence_transformer.modules.pooling.Pooling"
  }
]
  {
    "architectures": [
      "BertForMaskedLM"
    ],
    "attention_probs_dropout_prob": 0.1,
    "classifier_dropout": null,
    "hidden_act": "gelu",
    "hidden_dropout_prob": 0.1,
    "hidden_size": 768,
    "initializer_range": 0.02,
    "intermediate_size": 3072,
    "layer_norm_eps": 1e-12,
    "max_position_embeddings": 512,
    "model_type": "bert",
    "num_attention_heads": 12,
    "num_hidden_layers": 12,
    "pad_token_id": 0,
    "position_embedding_type": "absolute",
    "torch_dtype": "float32",
    "transformers_version": "4.12.5",
    "type_vocab_size": 2,
    "use_cache": true,
    "vocab_size": 6161
  }

Pre-training parameters used:

  {
    "epochs": 19, 
    "eval_steps": 10000, 
    "batch_size": 512, 
    "gradient_accumulation_steps": 1, 
    "mlm_prob": 0.4, 
    "bf16": True, 
    "tf32": True
  }

Output

The model outputs embeddings of size 768.

Data

The dataset is derived from the official Debian Repository. We used the distribution's build and packaging system apt to cross-compile source packages for multiple ISAs and optimisation levels. We used the standard optimisation levels supported by the distribution's compiler, O0, O1, O2, O3 and Os. We also added an optimization level that includes vectorization and cryptographic instructions that we refer to as Optimized-custom, or Oc. The binary files from each source package are disassembled using Ghidra and the resulting control-flow graphs (CFGs) are exported.

This results in a maximum of 24 (6 optimisation levels and 4 ISAs) potential CFGs which are semantically identical, i.e. they represent the same functionality. However, because not all source packages can be built for all ISAs and because some optimisation levels result in identical code, the actual number is often lower than that.

The dataset is split into a train, test and an evaluation set. This in done on source package basis, so all binaries and functions belonging to one source package are part of either the training set or the test set, not both.

Total amount of functions per architecture

Architecture # functions
amd64 8 202 164
i386 4 868 531
arm64 4 421 768
riscv64 3 791 434

Amount of functions per architecture per optimization

Architecture Optimization # functions
amd64 O0 1762442
amd64 O1 1448046
amd64 O2 1403180
amd64 O3 1308199
amd64 Oc 899892
amd64 Os 1380405
arm64 O0 918402
arm64 O1 824291
arm64 O2 790679
arm64 O3 741057
arm64 Oc 331736
arm64 Os 815603
i386 O0 1081475
i386 O1 873128
i386 O2 845943
i386 O3 816186
i386 Oc 412503
i386 Os 839296
riscv64 O0 845154
riscv64 O1 689670
riscv64 O2 670651
riscv64 O3 641428
riscv64 Oc 259691
riscv64 Os 684840

Preprocessing

Several preprocessing steps were taken. The CFGs are processed into a flat token list using ISA aware tokenizers. Immediate operands and addresses have been normalized to the nearest power-of-two to reduce token explosion caused by raw numeric values. Finally, target addresses of jumps, conditional branches or call instructions are normalised to function-relative JUMP_ADDR_* tokens.

Performance

We report two performance metrics: Mean Reciprocal Rank (MRR) and Accuracy@1.

These metrics are suitable for evaluating tasks where a single example is retrieved from a large collection. Accuracy@1 shows how many times the positive example ranked first. It is a useful metric, but it's also very strict, which is why we used it combined with MRR.

model ARM64 mrr ARM64 acc AMD64 mrr AMD64 acc RISC-V mrr RISC-V acc I386 mrr I386 acc crosslingual mrr crosslingual acc
Multilingual-ASMBERT 0.762 0.685 0.848 0.803 0.615 0.514 0.660 0.575 0.631 0.530
Downloads last month
69
Safetensors
Model size
90.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for NetherlandsForensicInstitute/Multilingual-ASMBERT