#

ai-interpretability

Here are 3 public repositories matching this topic...

Wondermongering / LinguisticPerturber

Probing linguistic robustness in transformers: a quantum-inspired approach to AI interpretability

machine-learning natural-language-processing word-embeddings computational-linguistics ai-safety probabilistic-models adversarial-examples perturbation-analysis transformer-models ai-interpretability language-model-analysis

Updated Mar 2, 2025
Python

kou-saki / i-asked-it-to-forget

I Asked It to Forget, but It Didn't — A Case of Miscommunication Between AI and Humans

Updated Apr 17, 2025

AlexTMjugador / redwoodresearch-interp-docker

📦 Redwood Research's transformer interpretability tools, conveniently packaged in a Docker container for simple and reproducible deployments.

docker ai ai-safety redwood-research ai-interpretability

Updated Apr 21, 2024
Dockerfile

Improve this page

Add a description, image, and links to the ai-interpretability topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the ai-interpretability topic, visit your repo's landing page and select "manage topics."

pFad - Phonifier reborn

Pfad - The Proxy pFad of © 2024 Garber Painting. All rights reserved.

Note: This service is not intended for secure transactions such as banking, social media, email, or purchasing. Use at your own risk. We assume no liability whatsoever for broken pages.

Alternative Proxies:

Alternative Proxy