SentenceTransformer based on google/embeddinggemma-300m
This is a
sentence-transformers
model finetuned from
google/embeddinggemma-300m
. It maps sentences & paragraphs to a 768-dimensional dense vector space and can be used for semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and more.
from sentence_transformers import SentenceTransformer
# Download from the 🤗 Hub
model = SentenceTransformer("yasserrmd/geo-gemma-300m-emb")
# Run inference
queries = [
"Based on the Brine Shrimp Lethality Test (BSLT), what are the toxicity levels of liquid smoke from cocoa pod skin at various pyrolysis temperatures and water contents?",
]
documents = [
'The Brine Shrimp Lethality Test (BSLT) was used to determine the toxicity levels of liquid smoke from cocoa pod skin at various pyrolysis temperatures and water contents. The results showed that the LC50 values (the concentration required to kill 50% of the test organisms) were as follows: at 200°C and 10% water content, 11,858.58 ppm; at 200°C and 15% water content, 13,094.23 ppm; at 200°C and 20% water content, 13,373.94 ppm; at 200°C and 25% water content, 15,703.52 ppm. At 300°C and 10% water content, 11,604.26 ppm; at 300°C and 15% water content, 11,673.05 ppm; at 300°C and 20% water content, 13,373.94 ppm; at 300°C and 25% water content, 13,373.94 ppm. At 400°C and 10% water content, 9,213.73 ppm; at 400°C and 15% water content, 13,094.237 ppm; at 400°C and 20% water content, 13,373.94 ppm; at 400°C and 25% water content, 12,493.63 ppm. All the results indicate that the liquid smoke from cocoa pod skin at different pyrolysis temperatures and water contents is classified as non-toxic.',
'The estimated annual flood damage for agriculture and built-up areas in the Tajan watershed, northern Iran, is projected to surge from USD 162 million to USD 376 million and USD 91 million to USD 220 million, respectively, by 2040, considering the land use change scenarios from 2021 to 2040.',
'The distribution of PM2.5 in Santa Ana, CA, tends to be higher in socioeconomically disadvantaged communities compared to other areas, highlighting environmental health inequities that persist in urban areas. This can inform policy decisions related to health equity and community access to resources.',
]
query_embeddings = model.encode_query(queries)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings.shape)
# [1, 768] [3, 768]# Get the similarity scores for the embeddings
similarities = model.similarity(query_embeddings, document_embeddings)
print(similarities)
# tensor([[0.5805, 0.0253, 0.0709]])
Training Details
Training Dataset
Unnamed Dataset
Size: 41,432 training samples
Columns:
sentence_0
and
sentence_1
Approximate statistics based on the first 1000 samples:
sentence_0
sentence_1
type
string
string
details
min: 12 tokens
mean: 27.1 tokens
max: 71 tokens
min: 17 tokens
mean: 119.32 tokens
max: 413 tokens
Samples:
sentence_0
sentence_1
How does plastic debris from land-based sources impact the ocean, particularly in the context of First Long Beach, China?
Plastic debris from land-based sources can significantly impact the ocean, as seen in the study conducted at First Long Beach (FLB), China. The study found that plastic debris amounts ranged from 2 to 82 particles per square meter on this marine sand beach. The most common size of plastics was 0.5–2.5 cm (44.4%), and the most common color was white (60.9%). The most abundant shape of plastic debris was fragments (76.2%). The amount of plastic debris varied significantly between different transects along the land-based source input zone due to the impacts of wind, ocean currents, and waves. Land-based wastewater discharge was identified as a major source of plastic debris on FLB, influenced by coastal water tide variations. Reduction strategies should focus on tracing and managing these land-based sources to mitigate the impact of plastic debris on the ocean.
How does the concentration of SO2 in urban areas of Nanjing correlate with the normalized difference vegetation index (NDVI), and what does this imply for public health?
The concentration of SO2 in urban areas of Nanjing exhibits a strong correlation (coefficient of determination, R2 > 0.5) with the normalized difference vegetation index (NDVI) within a radial distance of 2 km from the air pollutant monitoring sites. This indicates that NDVI can be an effective indicator for assessing the distribution and concentrations of air pollutants such as SO2. Negative correlations between NDVI and socio-economic indicators are observed under relatively consistent natural conditions, including climate and terrain. Therefore, the spatiotemporal distribution patterns of NDVI can provide valuable insights not only into socio-economic growth but also into the levels and locations of air pollution concentrations, which is crucial for public health interventions and policies.
How has the rise of user-generated geodata impacted the role of traditional map producers?
The rise of user-generated geodata has transformed ordinary citizens into neogeographers, blurring the boundaries between traditional map producers, such as national mapping agencies and local authorities, and citizens as consumers of this information. Citizens now actively participate in mapping different types of features on the Earth’s surface as volunteers, either by providing observations on the ground or tracing data from other sources, such as aerial photographs or satellite imagery. This has resulted in a significant increase in the availability of rich spatial datasets, which are often openly accessible through platforms like OpenStreetMap (OSM) and Ushahidi.
@inproceedings{reimers-2019-sentence-bert,
title = "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks",
author = "Reimers, Nils and Gurevych, Iryna",
booktitle = "Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing",
month = "11",
year = "2019",
publisher = "Association for Computational Linguistics",
url = "https://arxiv.org/abs/1908.10084",
}
MultipleNegativesRankingLoss
@misc{henderson2017efficient,
title={Efficient Natural Language Response Suggestion for Smart Reply},
author={Matthew Henderson and Rami Al-Rfou and Brian Strope and Yun-hsuan Sung and Laszlo Lukacs and Ruiqi Guo and Sanjiv Kumar and Balint Miklos and Ray Kurzweil},
year={2017},
eprint={1705.00652},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
Runs of yasserrmd geo-gemma-300m-emb on huggingface.co
83
Total runs
0
24-hour runs
0
3-day runs
-3
7-day runs
5
30-day runs
More Information About geo-gemma-300m-emb huggingface.co Model
geo-gemma-300m-emb huggingface.co
geo-gemma-300m-emb huggingface.co is an AI model on huggingface.co that provides geo-gemma-300m-emb's model effect (), which can be used instantly with this yasserrmd geo-gemma-300m-emb model. huggingface.co supports a free trial of the geo-gemma-300m-emb model, and also provides paid use of the geo-gemma-300m-emb. Support call geo-gemma-300m-emb model through api, including Node.js, Python, http.
geo-gemma-300m-emb huggingface.co is an online trial and call api platform, which integrates geo-gemma-300m-emb's modeling effects, including api services, and provides a free online trial of geo-gemma-300m-emb, you can try geo-gemma-300m-emb online for free by clicking the link below.
yasserrmd geo-gemma-300m-emb online free url in huggingface.co:
geo-gemma-300m-emb is an open source model from GitHub that offers a free installation service, and any user can find geo-gemma-300m-emb on GitHub to install. At the same time, huggingface.co provides the effect of geo-gemma-300m-emb install, users can directly use geo-gemma-300m-emb installed effect in huggingface.co for debugging and trial. It also supports api for free installation.