Explainability of CNN

Introduction # Capire in che modo una rete convoluzionale ci può dare insight migliori su come funzionano questi networks. Visualizzazione dei hidden layers # Slide visualization Potremmo fissare una immagine anche a caso, e modificare la x in modo che sia più simile a quanto…

Reading Time: 4 minutes · By Xuanqiang Angelo Huang

Expressiveness of NN

The perceptron # Slide summary of working of perceptron Note on the bias : it is only useful to move the treshhold where to consider the output to be 1 and where to be 1. Now we ask what can be predicted by a perceptron? We can see the update rule of the perceptron: { w = w + α…

Reading Time: 3 minutes · By Xuanqiang Angelo Huang

Object Detection

Introduction # Semantic segmentation # Vorremo trovare regioni che corrispondano a categorie diverse . E dividere in questo modo l’immagine secondo zone di informazione. Object detection # Vogliamo trovare il più piccolo box che vada a contenere l’oggetto. Questo è fatto con il…

Reading Time: 2 minutes · By Xuanqiang Angelo Huang

Ad-hoc Teamwork

Ad-Hoc Teamwork (AHT) in Reinforcement Learning # Problem Setting & Motivation # Ad-hoc teamwork concerns agents that must cooperate effectively with previously unknown teammates without any prior coordination, communication protocol, or shared learning history . This is…

Reading Time: 9 minutes · By Xuanqiang Angelo Huang

Distributional Reinforcement Learning

Distributional Reinforcement Learning # Motivation: Why Bother With the Whole Distribution? # Standard value-based RL collapses the random return into a single scalar via expectation: Q ( s , a ) = E [ Z ( s , a )] . The distributional perspective (Bellemare, Dabney, Munos,…

Reading Time: 15 minutes · By Xuanqiang Angelo Huang

Proximal Polixy Optimization

This document is DEPRECATED, please see RL Function Approximation . This documents attempts to briefly present the algorithm and some experiments found online about it. The following repo seems to be a good resource: here . Usually, PPO is explained as an actor critic framework…

Reading Time: 1 minutes · By Xuanqiang Angelo Huang

RL Losses

SDPO # See (Hübotter et al. 2026) GRPO # https://hlfshell.ai/posts/grpo/ GRPO (Group Relative Policy Optimization) comes from the DeepSeekMath paper. Its whole reason for existing is to get rid of the value/critic network that PPO needs. Instead of learning a separate model to…

Reading Time: 5 minutes · By Xuanqiang Angelo Huang

Introduction to Distributed Systems

Distributed computing addresses algorithms for a set of processes that seek to achieve some form of cooperation But it's quite a specific form of cooperation! The main failure mode was failure of operation of a single note, but in the modern case, we have different kinds of…

Reading Time: 1 minutes · By Xuanqiang Angelo Huang

Communication Games

AI Use Statement This note was fully AI generated as far as I recall, I don't remember which model. I noticed some people found this on the web. So take this with a grain of salt, I usually like to use AI generated summaries for personal study on the vault and don't publish…

Reading Time: 11 minutes · By Xuanqiang Angelo Huang

Information Bottleneck

These notes cover the core concepts of the Information Bottleneck method, widely used in machine learning and theoretical neuroscience. We start by defining the fundamental tension in learning and representation. Learning Design Goals # Compression : The representation should be…

Reading Time: 5 minutes · By Xuanqiang Angelo Huang