Instructions to use SebastianKotstein/restberta-qa-pm-ed with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SebastianKotstein/restberta-qa-pm-ed with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("question-answering", model="SebastianKotstein/restberta-qa-pm-ed")# Load model directly from transformers import AutoTokenizer, AutoModelForQuestionAnswering tokenizer = AutoTokenizer.from_pretrained("SebastianKotstein/restberta-qa-pm-ed") model = AutoModelForQuestionAnswering.from_pretrained("SebastianKotstein/restberta-qa-pm-ed", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 4,344 Bytes
51d4d96 26581b2 51d4d96 26581b2 bfcbf8a 5b28844 26581b2 5f33695 26581b2 dcbca5c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 | ---
license: cc-by-4.0
widget:
- text: The token for accessing this Web API
context: >-
auth.key location.city location.city_id location.country location.lat
location.lon location.postal_code state units
example_title: Weather API
- text: Returns an access token
context: >-
auth.post users.get users.post users.{userId}.address.get users.{userId}.address.put users.{userId}.delete users.{userId}.get users.{userId}.put
example_title: Test API
---
# RESTBERTa
RESTBERTa stands for *Representational State Transfer on Bidirectional Encoder Representations from Transformers approach* and should support machines
in processing structured syntax and unstructured natural language descriptions for semantics in Web API documentation.
In detail, we use question answering to solve the generic task of identifying a Web API syntax element (answer) in a syntax structure (paragraph) that matches the semantics described in a natural language query (question).
The identification and extraction of Web API syntax elements from Web API documentation is a common sub task of many Web API integration tasks, like parameter matching and endpoint discovery.
Thus, RESTBERTa might be a foundation for several Web API integration tasks.
Technically, RESTBERTa covers the concepts for fine-tuning a Transformer Encoder model, i.e., a pre-trained BERT model, to question answering with task-specific samples in order to prepare
a model for a specific Web API integration task.
The paper ["RESTBERTa: a Transformer-based question answering approach for semantic search in Web API documentation"](https://link.springer.com/article/10.1007/s10586-023-04237-x) demonstrates the application of RESTBERTa
to semantic parameter matching and endpoint discovery:
# RESTBERTa for Parameter Matching and Endpoint Discovery:
This repository contains the weights of a CodeBERT base model that has been fine-tuned to the task of parameter matching and endpoint discovery. For this, we formulate question answering as a multiple choice task:
Given a query in natural language that describes the purpose and behavior of the target parameter or endpoint, i.e., its semantics, the model should choose the parameter/endpoint from
a hierarchical structure of parameters or endpoints, i.e., a schema or a URI model.
Note: BERT models are optimized for linear text input. We, therefore, serialize schemas and URI models into linear text
by converting parameters and endpoints into an XPath-like notation, e.g., "users.{userId}.get" for an endpoint "GET /users/{userId}". The result is a list of alphabetically
sorted XPaths, e.g., "users.get users.post users.{userId}.address.get users.{userId}.address.put users.{userId}.delete users.{userId}.get users.{userId}.put".
# Fine-tuning
We fine-tuned the pre-trained microsoft/codebert-base model to the downstream task of question answering with 909,089 question-answering samples from 2,321
real-world OpenAPI documentation. Each sample consists of:
- Question: The natural language description of the parameter or endpoint, e.g., "Creates a new user"
- Answer: The parameter or endpoint in an XPath-like notation, e.g., "users.post"
- Paragraph: The hierarchical structure containing the parameter or endpoint, which is a list of parameters/endpoints in XPath-like notation, e.g., "users.get users.post users.{userId}.address.get users.{userId}.address.put users.{userId}.delete users.{userId}.get users.{userId}.put".
# Inference:
RESTBERTa requires a special output interpreter that processes the predictions made by the model in order to determine the suggested parameter or endpoint. We discuss the details in the paper.
# Hyperparameters:
The model was fine-tuned with ten epochs and a batch size of 16 on an Nvidia Ampere GPU. This repository contains the model checkpoint (weights) after five epochs of fine-tuning, which achieved the highest accuracy applied to our parameter-matching validation set.
# Citation:
```bibtex
@ARTICLE{10.1007/s10586-023-04237-x0,
author={Kotstein, Sebastian and Decker, Christian},
journal={Cluster Computing},
title={RESTBERTa: a Transformer-based question answering approach for semantic search in Web API documentation},
year={2024},
volume={},
number={},
pages={},
publisher={Springer}
doi={10.1007/s10586-023-04237-x}}
``` |