Three papers at NeurIPS 2026

This year our group has three papers at NeurIPS 2026 — one in the main conference and two at workshops — and together they tell a fairly coherent story about what it takes to make AI useful in dermatology. CD-RCM (main conference) tackles the imaging side: reflectance confocal microscopy gives clinicians cellular-resolution views of living skin, but only as sparse, anisotropic depth stacks, and we show how a feedforward transformer can fill in the missing depths in under a second. Semantically Coherent Calibration (Med-Reasoner workshop) tackles trust: a vision-language model can look well-calibrated on average while quietly over- or under-stating malignancy risk for specific patient subgroups, and we show how to discover those subgroups in a form clinicians can actually read. Grounding MLLMs with Quantitative Skin Attributes (VLM4RWD workshop) tackles interpretability: we fine-tune a multimodal LLM to predict measurable lesion attributes and find that the resulting embeddings are both clinically steerable and, without any diagnostic supervision, more diagnostic than existing dermatology foundation models. Seeing more, knowing when to trust the model, and explaining its reasoning in clinical terms — that is the thread running through all three.

CD-RCM: Generalizable Continuous-Depth Novel View Synthesis for Reflectance Confocal Microscopy

Tooba Imtiaz, Milind Rajadhyaksha, Kivanc Kose, Jennifer Dy
Main conferenceNeurIPS 2026 poster · OpenReview · Presentation

Reflectance confocal microscopy (RCM) provides noninvasive “optical biopsies” of skin by capturing en-face images at successive depths. The catch is that lateral resolution (~0.5 µm) is about six times finer than axial resolution (~3 µm), so the resulting z-stacks are strongly anisotropic — fine for viewing slice by slice, but poor for the histopathology-style cross-sections clinicians are trained on. Spline interpolation between slices produces blocky, anatomically implausible transitions.

CD-RCM is the first novel-view synthesis method built specifically for RCM. The key observation is that RCM’s acquisition geometry is nothing like a camera orbiting an object: there is no parallax, only pure axial translation through tissue. We model the microscope as a virtual pinhole camera translating along z, which lets us use Plücker ray embeddings and a decoder-only transformer (building on LVSM) to synthesize slices at arbitrary query depths from just three input slices — in a single feedforward pass, with no per-stack optimization. To preserve cellular texture, we also introduce a skin-specific perceptual loss built on a DINOv3 backbone adapted to RCM data via LoRA.

Overview of the CD-RCM architecture
Overview of CD-RCM. Sparse input slices and their Plücker ray embeddings are tokenized and processed by a decoder-only transformer; target ray tokens condition the synthesis of unseen depths. Training combines photometric, LPIPS, and skin-specific perceptual losses.
Sagittal, coronal and oblique cross-sections of input and CD-RCM-densified stacks
Cross-sectional and arbitrary-plane views of densified stacks. CD-RCM removes the staircase artifacts visible in sparsely sampled input stacks, revealing continuous tissue morphology.

Learning Semantically Coherent Calibration Groups

Brandon Dominique, Prudence Lam, Max Torop, Nicholas Kurtansky, Jochen Weber, Kivanc Kose, Veronica Rotemberg, Jennifer Dy
WorkshopNeurIPS 2026 Med-Reasoner poster · OpenReview

A model’s confidence shapes whether a clinician trusts or overrides it. Standard calibration methods fix confidence in aggregate, and multicalibration extends this to predefined groups — but in practice the groups that matter are often unknown. Existing group-discovery methods either optimize purely for calibration (and collapse into one or two uninterpretable groups) or purely for semantic similarity (and produce coherent-looking groups that still mix well- and poorly-calibrated samples).

Semantically Coherent Calibration (SCC) jointly optimizes both. A learned soft grouping function partitions a VLM’s multimodal embedding space so that samples in each group share both a calibration need and a semantic neighborhood, and each group gets its own Platt or temperature scaling parameters. The result is a set of groups a practitioner can inspect: on ISIC 2024, for example, the group requiring the largest downward correction skews older and 75% male — a concrete, testable hypothesis about where the model’s training data may be thin.

Illustration of group discovery along calibration-need and semantic axes
Group discovery along two independent axes: calibration need (color) and semantic category (shape). (a) Grouping solely on calibration need mixes distinct semantic categories together. (b) Grouping solely on semantic similarity produces semantically coherent groups that still mix well- and poorly-calibrated samples, leaving calibration need unaddressed within each group. (c) SCC jointly optimizes both objectives, discovering groups that are homogeneous in calibration need and semantic category, with interpretable descriptions (e.g., “Mostly male, older, lower extremity”).
Representative images from the four groups SCC discovers on ISIC 2024
Representative images from the four groups SCC discovers on ISIC 2024, annotated by risk level and shared characteristics (a), compared with the groups found by the GC+TS baseline (b).

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

Max Torop, Masih Eskandar, Nicholas Kurtansky, Jinyang Liu, Jochen Weber, Octavia Camps, Veronica Rotemberg, Jennifer Dy, Kivanc Kose
WorkshopNeurIPS 2026 VLM4RWD poster · OpenReview · Presentation

Skin cancer classifiers can match dermatologists on accuracy while learning to rely on rulers, hair, and surgical markings rather than the lesion. Multimodal LLMs promise interpretable, conversational reasoning, but they give little control over which visual information their representations actually encode — and they are notoriously bad at quantification.

We fine-tune Qwen2-VL on the SLICE-3D dataset of 3D total-body-photography tiles to predict 16 quantitative lesion attributes known to be predictive of malignancy: area, diameter, border irregularity, color asymmetry, lesion–skin contrast, and more. Because the decoder merges image and text, the same image embedding can be steered at query time toward any attribute — or any combination of attributes — simply by asking. This turns a standard embedding into a clinically transparent vector for composed retrieval (“find lesions like this one, but matched on area and border irregularity”), with a hierarchical retrieval scheme that avoids storing an embedding per attribute combination.

Image-only and attribute-conditioned embedding functions in the MLLM decoder
Image-only vs. attribute-conditioned embeddings. The image embedding averages penultimate-layer image tokens; the attribute-conditioned embedding is the final token after appending a question such as "What is the area in mm²?"
Top-5 retrieval results for image-only and area-conditioned embeddings
Top-5 retrieval for three query lesions. Image-only retrieval returns visually similar lesions; area-conditioned retrieval returns lesions that are visually similar and matched on area.

← Back to News