RayBench EmbeddedInteractive engineering labs
ARCHITECTURE

Cortex-M

Arm Cortex-M programmers model, Thumb-2 ISA, memory-mapped peripherals, and NVIC interrupt handling. Interactive register and stack simulator.

Reviewed 2026-08-224,232 words

Arm Cortex-M: The One You Will Actually Ship

Cortex-M is the default architecture for embedded work, and if a reader learns exactly one processor properly this should be it. Its design goal was not peak throughput but deterministic, low-latency response with a programming model that ordinary C can reach - which is why its most distinctive features are all about exceptions rather than arithmetic. The hardware stacks registers for you on interrupt entry, which means a plain C function can be an interrupt handler with no assembly wrapper at all. Coming to it after the 8051, the six programmer's-model questions have familiar shapes and richer answers, and the differences are the lesson: a single flat address space instead of four, load/store instead of memory operands, two stack pointers instead of one, and an interrupt controller that does most of the work a 8051 handler had to do by hand.

How it is built

  • Sixteen registers, of which R0 to R12 are general purpose and the last three are structural: R13 is the stack pointer, R14 the link register holding the return address, R15 the program counter. The link register is the significant one - a function call does not push a return address to the stack, it puts it in LR, so a leaf function that calls nothing needs no stack traffic at all to return.
  • There are two stack pointers, MSP and PSP, and R13 is whichever one is currently selected. Handler mode always uses MSP; Thread mode can use either. That split is what lets an RTOS give each task its own stack while interrupts run on a single known one, and it is why a stack overflow in a task does not necessarily corrupt the kernel.
  • The instruction set is Thumb-2: a mix of 16-bit and 32-bit encodings in one stream, chosen per instruction for density. There is no separate 32-bit mode to switch into - Cortex-M executes only Thumb, which is why the vector table's addresses have their low bit set and why clearing it faults immediately.
  • It is a load/store architecture. Data moves between registers and memory only through LDR and STR and their variants; everything else is register to register. The APSR carries N, Z, C and V, and whether an instruction updates them is part of the encoding - ADD leaves flags alone, ADDS sets them, and forgetting the S is a common source of branches that never fire.
  • Exception entry is the architecture's centrepiece. On an interrupt the hardware pushes eight registers - R0-R3, R12, LR, PC and xPSR - before the handler runs, so the handler may use those freely and the AAPCS calling convention is satisfied. That is what makes a C function directly usable as an ISR, and it is the single biggest practical difference from the 8051.
  • The exception mechanism has optimisations that matter for latency. Tail-chaining skips the unstack-then-restack between two back-to-back interrupts; late arrival lets a higher-priority interrupt take over an entry sequence already in progress. Both reduce worst-case latency, and both are architectural guarantees rather than implementation luck.
  • Memory is one flat 4 GB space with an architecturally defined layout: code low, SRAM next, peripherals above that, and the System Control Space near the top holding the NVIC and SCB. There is no MMU. An optional MPU can enforce region permissions but does not translate addresses, so a pointer means the same thing everywhere and there is no per-process address space.
  • The vector table is the boot contract. Word zero is the initial stack pointer value, word one is the reset handler's address, and the rest are exception handlers. The processor loads SP from the table before executing a single instruction, which is why a corrupt first word produces a fault before main can possibly run.

Design procedure

  1. Answer the six questions again on Cortex-M and diff them against your 8051 answers. The differences are the content of this topic.
  2. Read the vector table in a real startup file and identify the first two words. Understanding that the hardware loads SP from word zero explains most early boot failures.
  3. Write an ISR as a plain C function and check the disassembly for the absence of a save-and-restore preamble. That absence is the automatic stacking doing its job.
  4. Find an ADDS in compiled output and work out why the compiler chose the flag-setting form. Flags-as-encoding is the detail that separates reading Thumb from guessing at it.
  5. Look at where MSP and PSP are selected in an RTOS port. The two-stack split is invisible in bare-metal code and central to everything above it.
  6. Then go to the NVIC topic. Registers and instructions are half of Cortex-M; the exception controller is the half that decides whether firmware meets its deadlines.

Key terms

Thumb-2
Mixed 16- and 32-bit encodings in one stream. Cortex-M executes only Thumb, which is why vector addresses have bit 0 set.
LR (R14)
The link register holding the return address. A leaf function returns with no stack traffic at all.
MSP / PSP
Two stack pointers. Handler mode always uses MSP; an RTOS gives tasks their own PSP.
Automatic stacking
The hardware pushes R0-R3, R12, LR, PC and xPSR on entry, so a C function can be an ISR unmodified.
EXC_RETURN
The magic LR value on exception entry. Returning to it, rather than to a normal address, is what unstacks.
Tail-chaining
Back-to-back interrupts skip the unstack and restack between them. An architectural latency guarantee.
S suffix
Whether an instruction updates flags is part of its encoding. ADD does not; ADDS does.
Vector table
Word 0 is the initial SP, word 1 the reset handler. The processor reads SP before executing anything.

Worked example

@ Cortex-M against the 8051, on the same job: toggle a pin in an ISR.

        .section .isr_vector
        .word   _estack             @ word 0: hardware loads SP from here
        .word   Reset_Handler + 1   @ word 1: +1 because Thumb. Clearing
                                    @ that bit faults on the first fetch.

@ A plain C function works as an ISR. No PUSH wall, because the
@ hardware already stacked R0-R3, R12, LR, PC and xPSR for us.
TIM2_IRQHandler:
        ldr     r0, =GPIOA_ODR      @ load/store: memory only via LDR/STR
        ldr     r1, [r0]
        eor     r1, r1, #1
        str     r1, [r0]
        bx      lr                  @ lr holds EXC_RETURN, so this unstacks

# What the hardware stacks, and why those eight:
#
#   R0 R1 R2 R3 R12   the AAPCS caller-saved set
#   LR                so the handler can call other functions
#   PC                where to resume
#   xPSR              flags and the interrupted state
#
#   Exactly the registers a C function is allowed to clobber. That is
#   not a coincidence - it is why no assembly wrapper is needed.

# The 8051 comparison, kept from the previous topic:
#
#                        8051              Cortex-M
#   hardware saves       PC only           8 registers
#   handler prologue     PUSH ACC, PSW...  none
#   address spaces       four              one flat
#   memory operands      yes               no - load/store
#   stack pointers       one               two (MSP, PSP)

Common pitfalls

General Data Processing & I/O Control

Basic arithmetic, logic, and control instructions available on all Cortex-M cores

ADC
Add with Carry - Rd = Rn + Rm + Carry - ADC{S}{cond} Rd, Rn, Rm
ADD
Add - Rd = Rn + operand2 - ADD{S}{cond} Rd, Rn, #imm | Rm
AND
Bitwise AND - Rd = Rn & Rm - AND{S}{cond} Rd, Rn, Rm
ASR
Arithmetic Shift Right - Rd = Rn >> imm (sign-extended) - ASR{S}{cond} Rd, Rn, #imm | Rs
B
Branch - PC = label - B{cond} label
BIC
Bit Clear - Rd = Rn & ~Rm - BIC{S}{cond} Rd, Rn, Rm
BKPT
Breakpoint - Enter debug state - BKPT #imm
BL
Branch with Link - LR = PC+4; PC = label - BL label
BLX
Branch with Link and Exchange - LR = PC+4; PC = Rm; T bit = Rm[0] - BLX Rm | label
BX
Branch and Exchange - PC = Rm; T bit = Rm[0] - BX Rm
CBNZ
Compare and Branch on Non-Zero - if (Rn != 0) PC = label - CBNZ Rn, label
CBZ
Compare and Branch on Zero - if (Rn == 0) PC = label - CBZ Rn, label
CLZ
Count Leading Zeros - Rd = number of leading zeros in Rn - CLZ Rd, Rn
CMN
Compare Negative - Flags = Rn + Rm - CMN{cond} Rn, Rm
CMP
Compare - Flags = Rn - operand2 - CMP{cond} Rn, Rm | #imm
CPS
Change Processor State - Set/clear I or F interrupt masks - CPSID i | f | CPSIE i | f
DMB
Data Memory Barrier - Wait for all memory accesses to complete - DMB {option}
DSB
Data Synchronization Barrier - Complete all memory accesses - DSB {option}
EOR
Exclusive OR - Rd = Rn ^ Rm - EOR{S}{cond} Rd, Rn, Rm
ISB
Instruction Synchronization Barrier - Flush pipeline and refetch instructions - ISB {option}
LDMIA
Load Multiple Increment After - Load registers from sequential memory - LDMIA Rn!, {Rlist}
LDR
Load Register - Rd = Memory[Rn + offset] - LDR{cond} Rd, [Rn, #offset]
LDRB
Load Register Byte - Rd = zero_extend(Memory[Rn + offset]) - LDRB{cond} Rd, [Rn, #offset]
LDRH
Load Register Halfword - Rd = zero_extend(Memory[Rn + offset]) - LDRH{cond} Rd, [Rn, #offset]
LDRSB
Load Register Signed Byte - Rd = sign_extend(Memory[Rn + offset]) - LDRSB{cond} Rd, [Rn, #offset]
LDRSH
Load Register Signed Halfword - Rd = sign_extend(Memory[Rn + offset]) - LDRSH{cond} Rd, [Rn, #offset]
LSL
Logical Shift Left - Rd = Rn << imm - LSL{S}{cond} Rd, Rn, #imm | Rs
LSR
Logical Shift Right - Rd = Rn >> imm (zero-filled) - LSR{S}{cond} Rd, Rn, #imm | Rs
MOV
Move - Rd = operand2 - MOV{S}{cond} Rd, #imm | Rm
MUL
Multiply - Rd = Rn * Rm (lower 32 bits) - MUL{S}{cond} Rd, Rn, Rm
MVN
Move Not - Rd = ~Rm - MVN{S}{cond} Rd, Rm
NOP
No Operation - None - NOP
ORR
Bitwise OR - Rd = Rn | Rm - ORR{S}{cond} Rd, Rn, Rm
POP
Pop Registers from Stack - Load registers from [SP]; SP += 4*n - POP {Rlist}
PUSH
Push Registers to Stack - SP -= 4*n; Store registers to [SP] - PUSH {Rlist}
REV
Reverse Byte Order - Rd = byte_reverse(Rn) - REV Rd, Rn
REV16
Reverse Byte Order in Halfwords - Reverse bytes in each 16-bit half - REV16 Rd, Rn
REVSH
Reverse Signed Halfword - Reverse bytes in low halfword, sign-extend - REVSH Rd, Rn
ROR
Rotate Right - Rd = Rn rotated right by Rm - ROR{S}{cond} Rd, Rn, Rm
RSB
Reverse Subtract - Rd = Rm - Rn - RSB{S}{cond} Rd, Rn, Rm
SDIV
Signed Divide - Rd = Rn / Rm (signed) - SDIV{cond} Rd, Rn, Rm
STMIA
Store Multiple Increment After - Store registers to sequential memory - STMIA Rn!, {Rlist}
STR
Store Register - Memory[Rn + offset] = Rd - STR{cond} Rd, [Rn, #offset]
STRB
Store Register Byte - Memory[Rn + offset] = Rd[7:0] - STRB{cond} Rd, [Rn, #offset]
STRH
Store Register Halfword - Memory[Rn + offset] = Rd[15:0] - STRH{cond} Rd, [Rn, #offset]
SUB
Subtract - Rd = Rn - operand2 - SUB{S}{cond} Rd, Rn, #imm | Rm
SVC
Supervisor Call - Trigger SVC exception - SVC #imm
SXTB
Sign-Extend Byte - Rd = sign_extend(Rn[7:0]) - SXTB Rd, Rn
SXTH
Sign-Extend Halfword - Rd = sign_extend(Rn[15:0]) - SXTH Rd, Rn
TST
Test - Flags = Rn & Rm - TST{cond} Rn, Rm
UDIV
Unsigned Divide - Rd = Rn / Rm (unsigned) - UDIV{cond} Rd, Rn, Rm
UXTB
Zero-Extend Byte - Rd = zero_extend(Rn[7:0]) - UXTB Rd, Rn
UXTH
Zero-Extend Halfword - Rd = zero_extend(Rn[15:0]) - UXTH Rd, Rn
WFE
Wait For Event - Wait for event signal - WFE
WFI
Wait For Interrupt - Wait for interrupt - WFI
YIELD
Yield - Hint to processor for thread switching - YIELD

Advanced Data Processing & Bit Field

Enhanced instructions for complex data manipulation and bit field operations

ADR
Address to Register (PC-relative) - Rd = PC + offset to label - ADR Rd, label
BFC
Bit Field Clear - Clear width bits starting at lsb - BFC Rd, #lsb, #width
BFI
Bit Field Insert - Copy width bits from Rn to Rd at lsb - BFI Rd, Rn, #lsb, #width
CDP
Coprocessor Data Processing - Coprocessor-specific operation - CDP p#, opcode, CRd, CRn, CRm
DBG
Debug - Debug hint - DBG #imm
LDC
Load Coprocessor - Load coprocessor register from memory - LDC p#, CRd, [Rn, #offset]
LDM
Load Multiple - Load multiple registers - LDM{mode} Rn, {Rlist}
LDRD
Load Register Dual - Load 64-bit value to Rd:Rt - LDRD Rd, Rt, [Rn, #offset]
LDMDB
Load Multiple Decrement Before - Decrement then load - LDMDB Rn!, {Rlist}
MCR
Move to Coprocessor from Register - Copy Rd to coprocessor register - MCR p#, opcode, Rd, CRn, CRm
MCRR
Move to Coprocessor from Two Registers - Copy Rd:Rt to coprocessor - MCRR p#, opcode, Rd, Rt, CRm
MLA
Multiply Accumulate - Rd = (Rn * Rm) + Ra - MLA{S}{cond} Rd, Rn, Rm, Ra
MLS
Multiply Subtract - Rd = Ra - (Rn * Rm) - MLS{cond} Rd, Rn, Rm, Ra
MRC
Move to Register from Coprocessor - Copy coprocessor register to Rd - MRC p#, opcode, Rd, CRn, CRm
MRRC
Move to Two Registers from Coprocessor - Copy coprocessor to Rd:Rt - MRRC p#, opcode, Rd, Rt, CRm
MRS
Move to Register from Special Register - Rd = special register - MRS Rd, spec_reg
MSR
Move to Special Register from Register - special register = Rn - MSR spec_reg, Rn
PLD
Prefetch Data - Hint to prefetch memory - PLD [Rn, #offset]
PLI
Prefetch Instruction - Hint to prefetch instruction - PLI [Rn, #offset]
RBIT
Reverse Bits - Rd = bit_reverse(Rn) - RBIT Rd, Rn
SBFX
Signed Bit Field Extract - Extract width bits from lsb, sign-extend - SBFX Rd, Rn, #lsb, #width
SDIV
Signed Divide - Rd = Rn / Rm (signed) - SDIV{cond} Rd, Rn, Rm
SEL
Select Bytes - Select bytes from Rn or Rm based on GE - SEL Rd, Rn, Rm
SEV
Send Event - Generate event signal - SEV
SMC
Secure Monitor Call - Enter secure monitor mode - SMC #imm
SMMLA
Signed Most Significant Word Multiply Accumulate - Rd = Ra + (Rn * Rm)[63:32] - SMMLA{R} Rd, Rn, Rm, Ra
SMMLS
Signed Most Significant Word Multiply Subtract - Rd = Ra - (Rn * Rm)[63:32] - SMMLS{R} Rd, Rn, Rm, Ra
SMMUL
Signed Most Significant Word Multiply - Rd = (Rn * Rm)[63:32] - SMMUL{R} Rd, Rn, Rm
STM
Store Multiple - Store multiple registers - STM{mode} Rn, {Rlist}
STRD
Store Register Dual - Store 64-bit value from Rd:Rt - STRD Rd, Rt, [Rn, #offset]
STREX
Store Exclusive - Attempt exclusive store, Rd = success/fail - STREX Rd, Rn, [Rm]
STREXB
Store Exclusive Byte - Exclusive byte store - STREXB Rd, Rn, [Rm]
STREXH
Store Exclusive Halfword - Exclusive halfword store - STREXH Rd, Rn, [Rm]
STRT
Store Register Unprivileged - Store as unprivileged access - STRT Rd, [Rn]
SUBS
Subtract with flags - Rd = Rn - Rm; update flags - SUBS Rd, Rn, Rm
SXTAB
Sign-Extend and Add Byte - Rd = Rn + sign_extend(Rm[7:0]) - SXTAB Rd, Rn, Rm
SXTAB16
Sign-Extend and Add Two Halfwords - Sign-extend bytes to halfwords and add - SXTAB16 Rd, Rn, Rm
SXTAH
Sign-Extend and Add Halfword - Rd = Rn + sign_extend(Rm[15:0]) - SXTAH Rd, Rn, Rm
UBFX
Unsigned Bit Field Extract - Extract width bits from lsb, zero-extend - UBFX Rd, Rn, #lsb, #width
UDIV
Unsigned Divide - Rd = Rn / Rm (unsigned) - UDIV{cond} Rd, Rn, Rm
UXTAB
Zero-Extend and Add Byte - Rd = Rn + zero_extend(Rm[7:0]) - UXTAB Rd, Rn, Rm
UXTAB16
Zero-Extend and Add Two Halfwords - Zero-extend bytes to halfwords and add - UXTAB16 Rd, Rn, Rm
UXTAH
Zero-Extend and Add Halfword - Rd = Rn + zero_extend(Rm[15:0]) - UXTAH Rd, Rn, Rm

DSP Instructions (SIMD & Fast MAC)

Digital Signal Processing instructions with SIMD and Multiply-Accumulate operations

QADD
Saturating Add - Rd = saturate(Rn + Rm) - QADD{cond} Rd, Rn, Rm
QDADD
Saturating Double and Add - Rd = Rn + saturate(Rm * 2) - QDADD{cond} Rd, Rn, Rm
QDSUB
Saturating Double and Subtract - Rd = Rn - saturate(Rm * 2) - QDSUB{cond} Rd, Rn, Rm
QSUB
Saturating Subtract - Rd = saturate(Rn - Rm) - QSUB{cond} Rd, Rn, Rm
SADD16
Signed Add 16 - Rd[15:0] = Rn[15:0] + Rm[15:0]; Rd[31:16] = Rn[31:16] + Rm[31:16] - SADD16 Rd, Rn, Rm
SADD8
Signed Add 8 - Four parallel 8-bit signed additions - SADD8 Rd, Rn, Rm
SASX
Signed Add and Subtract with Exchange - Rd[31:16] = Rn[31:16] + Rm[15:0]; Rd[15:0] = Rn[15:0] - Rm[31:16] - SASX Rd, Rn, Rm
SHADD16
Signed Halving Add 16 - Rd[15:0] = (Rn[15:0] + Rm[15:0]) >> 1 - SHADD16 Rd, Rn, Rm
SHASX
Signed Halving Add and Subtract with Exchange - Halving add/subtract with exchange - SHASX Rd, Rn, Rm
SHSAX
Signed Halving Subtract and Add with Exchange - Halving subtract/add with exchange - SHSAX Rd, Rn, Rm
SHSUB16
Signed Halving Subtract 16 - Rd[15:0] = (Rn[15:0] - Rm[15:0]) >> 1 - SHSUB16 Rd, Rn, Rm
SHSUB8
Signed Halving Subtract 8 - Four parallel halving 8-bit subtractions - SHSUB8 Rd, Rn, Rm
SMLA
Signed Multiply Accumulate - Rd = Ra + (Rn * Rm) with saturation - SMLA{B|T}{cond} Rd, Rn, Rm, Ra
SMLAD
Signed Multiply Accumulate Dual - Rd = Ra + Rn[15:0]*Rm[15:0] + Rn[31:16]*Rm[31:16] - SMLAD{X}{cond} Rd, Rn, Rm, Ra
SMLAL
Signed Multiply Accumulate Long - 64-bit accumulate: RdHi:RdLo += Rn * Rm - SMLAL{cond} RdLo, RdHi, Rn, Rm
SMLALD
Signed Multiply Accumulate Long Dual - 64-bit dual multiply accumulate - SMLALD{X}{cond} RdLo, RdHi, Rn, Rm
SMLSD
Signed Multiply Subtract Dual - Rd = Ra + Rn[15:0]*Rm[15:0] - Rn[31:16]*Rm[31:16] - SMLSD{X}{cond} Rd, Rn, Rm, Ra
SMLSLD
Signed Multiply Subtract Long Dual - 64-bit dual multiply subtract - SMLSLD{X}{cond} RdLo, RdHi, Rn, Rm
SMMLA
Signed Most Significant Word Multiply Accumulate - Rd = Ra + (Rn * Rm)[63:32] - SMMLA{R}{cond} Rd, Rn, Rm, Ra
SMUAD
Signed Multiply Dual - Rd = Rn[15:0]*Rm[15:0] + Rn[31:16]*Rm[31:16] - SMUAD{X}{cond} Rd, Rn, Rm
SMUSD
Signed Multiply Subtract Dual - Rd = Rn[15:0]*Rm[15:0] - Rn[31:16]*Rm[31:16] - SMUSD{X}{cond} Rd, Rn, Rm
SSAT
Signed Saturate - Rd = saturate_signed(Rn, imm) - SSAT Rd, #imm, Rn
SSAT16
Signed Saturate 16 - Saturate both halfwords to signed imm bits - SSAT16 Rd, #imm, Rn
SSAX
Signed Subtract and Add with Exchange - Rd[31:16] = Rn[31:16] - Rm[15:0]; Rd[15:0] = Rn[15:0] + Rm[31:16] - SSAX Rd, Rn, Rm
SSUB16
Signed Subtract 16 - Two parallel 16-bit signed subtractions - SSUB16 Rd, Rn, Rm
SSUB8
Signed Subtract 8 - Four parallel 8-bit signed subtractions - SSUB8 Rd, Rn, Rm
UADD16
Unsigned Add 16 - Two parallel 16-bit unsigned additions - UADD16 Rd, Rn, Rm
UADD8
Unsigned Add 8 - Four parallel 8-bit unsigned additions - UADD8 Rd, Rn, Rm
UASX
Unsigned Add and Subtract with Exchange - Rd[31:16] = Rn[31:16] + Rm[15:0]; Rd[15:0] = Rn[15:0] - Rm[31:16] - UASX Rd, Rn, Rm
UHADD16
Unsigned Halving Add 16 - Rd[15:0] = (Rn[15:0] + Rm[15:0]) >> 1 - UHADD16 Rd, Rn, Rm
UHASX
Unsigned Halving Add and Subtract with Exchange - Halving add/sub with exchange - UHASX Rd, Rn, Rm
UHSAX
Unsigned Halving Subtract and Add with Exchange - Halving subtract/add with exchange - UHSAX Rd, Rn, Rm
UHSUB16
Unsigned Halving Subtract 16 - Rd[15:0] = (Rn[15:0] - Rm[15:0]) >> 1 - UHSUB16 Rd, Rn, Rm
UHSUB8
Unsigned Halving Subtract 8 - Four parallel halving 8-bit subtractions - UHSUB8 Rd, Rn, Rm
UMAAL
Unsigned Multiply Accumulate Accumulate Long - RdHi:RdLo = Rn + Rm + (Rn * Rm) - UMAAL RdLo, RdHi, Rn, Rm
UMLAL
Unsigned Multiply Accumulate Long - RdHi:RdLo += Rn * Rm (unsigned) - UMLAL{cond} RdLo, RdHi, Rn, Rm
UMULL
Unsigned Multiply Long - RdHi:RdLo = Rn * Rm (unsigned 64-bit) - UMULL{cond} RdLo, RdHi, Rn, Rm
UQADD16
Unsigned Saturating Add 16 - Two parallel saturating 16-bit additions - UQADD16 Rd, Rn, Rm
UQADD8
Unsigned Saturating Add 8 - Four parallel saturating 8-bit additions - UQADD8 Rd, Rn, Rm
UQASX
Unsigned Saturating Add and Subtract with Exchange - Saturating add/sub with exchange - UQASX Rd, Rn, Rm
UQSAX
Unsigned Saturating Subtract and Add with Exchange - Saturating subtract/add with exchange - UQSAX Rd, Rn, Rm
UQSUB16
Unsigned Saturating Subtract 16 - Two parallel saturating 16-bit subtractions - UQSUB16 Rd, Rn, Rm
UQSUB8
Unsigned Saturating Subtract 8 - Four parallel saturating 8-bit subtractions - UQSUB8 Rd, Rn, Rm
USAD8
Unsigned Sum of Absolute Differences 8 - Rd = sum |Rn[i] - Rm[i]| for i=0..3 - USAD8 Rd, Rn, Rm
USADA8
Unsigned Sum of Absolute Differences 8 Accumulate - Rd = Ra + sum |Rn[i] - Rm[i]| - USADA8 Rd, Rn, Rm, Ra
USAT
Unsigned Saturate - Rd = saturate_unsigned(Rn, imm) - USAT Rd, #imm, Rn
USAT16
Unsigned Saturate 16 - Saturate both halfwords to unsigned imm bits - USAT16 Rd, #imm, Rn
USAX
Unsigned Subtract and Add with Exchange - Rd[31:16] = Rn[31:16] - Rm[15:0]; Rd[15:0] = Rn[15:0] + Rm[31:16] - USAX Rd, Rn, Rm
USUB16
Unsigned Subtract 16 - Two parallel 16-bit unsigned subtractions - USUB16 Rd, Rn, Rm
USUB8
Unsigned Subtract 8 - Four parallel 8-bit unsigned subtractions - USUB8 Rd, Rn, Rm
UXTB16
Unsigned Extend Byte 16 - Zero-extend Rn[7:0] and Rn[23:16] to halfwords - UXTB16 Rd, Rn

Floating Point Instructions

Single-precision floating point operations (VFPv4-SP)

VABS
Floating-Point Absolute - Sd = |Sm| - VABS.F32 Sd, Sm
VADD
Floating-Point Add - Sd = Sn + Sm - VADD.F32 Sd, Sn, Sm
VCMP
Floating-Point Compare - Compare Sn with Sm - VCMP.F32 Sn, Sm
VCMP
Floating-Point Compare with Zero - Compare Sn with 0.0 - VCMP.F32 Sn, #0
VCVT
Floating-Point Convert - Convert float to int or int to float - VCVT.S32.F32 Sd, Sm | VCVT.F32.S32 Sd, Sm
VCVTR
Floating-Point Convert with Rounding - Convert with current rounding mode - VCVTR.S32.F32 Sd, Sm
VDIV
Floating-Point Divide - Sd = Sn / Sm - VDIV.F32 Sd, Sn, Sm
VLD1
Vector Load One Structure - Load from memory to vector register - VLD1.32 {Sd}, [Rn]
VLD2
Vector Load Two Structures - Load and deinterleave two values - VLD2.32 {Sd, Se}, [Rn]
VLD3
Vector Load Three Structures - Load and deinterleave three values - VLD3.32 {Sd, Se, Sf}, [Rn]
VLD4
Vector Load Four Structures - Load and deinterleave four values - VLD4.32 {Sd, Se, Sf, Sg}, [Rn]
VMOV
Vector Move - Move value to vector register - VMOV.F32 Sd, Sm | VMOV Sd, #imm
VMRS
Vector Move to Register from Special - Copy FPSCR to general register - VMRS Rd, FPSCR
VMSR
Vector Move to Special from Register - Copy general register to FPSCR - VMSR FPSCR, Rd
VMUL
Vector Multiply - Sd = Sn * Sm - VMUL.F32 Sd, Sn, Sm
VNEG
Vector Negate - Sd = -Sm - VNEG.F32 Sd, Sm
VNMLA
Vector Negative Multiply Accumulate - Sd = -(Sd + Sn * Sm) - VNMLA.F32 Sd, Sn, Sm
VNMLS
Vector Negative Multiply Subtract - Sd = -(Sd - Sn * Sm) - VNMLS.F32 Sd, Sn, Sm
VNMUL
Vector Negative Multiply - Sd = -(Sn * Sm) - VNMUL.F32 Sd, Sn, Sm
VPOP
Vector Pop - Pop from stack to vector registers - VPOP {Slist}
VPUSH
Vector Push - Push vector registers to stack - VPUSH {Slist}
VRADDHN
Vector Rounding Add Narrow - Add and narrow with rounding - VRADDHN.D16 Dd, Dn, Dm
VRECPE
Vector Reciprocal Estimate - Sd ~= 1/Sm (approximate) - VRECPE.F32 Sd, Sm
VRECPS
Vector Reciprocal Step - Sd = 2 - Sn * Sm - VRECPS.F32 Sd, Sn, Sm
VRINTA
Vector Round to Integral (Away) - Round away from zero - VRINTA.F32 Sd, Sm
VRINTM
Vector Round to Integral (Minus) - Round toward -infinity - VRINTM.F32 Sd, Sm
VRINTN
Vector Round to Integral (Nearest) - Round to nearest (ties to even) - VRINTN.F32 Sd, Sm
VRINTP
Vector Round to Integral (Plus) - Round toward +infinity - VRINTP.F32 Sd, Sm
VRINTR
Vector Round to Integral (with Rounding) - Round using FPSCR rounding mode - VRINTR.F32 Sd, Sm
VRINTX
Vector Round to Integral (Exact) - Round to integral, set IXF if inexact - VRINTX.F32 Sd, Sm
VRINTZ
Vector Round to Integral (Zero) - Round toward zero (truncate) - VRINTZ.F32 Sd, Sm
VRSHR
Vector Round Shift Right - Shift right with rounding - VRSHR.S32 Dd, Dm, #imm
VRSQRT
Vector Reciprocal Square Root Estimate - Sd ~= 1/sqrtSm - VRSQRT.F32 Sd, Sm
VRSQRTS
Vector Reciprocal Square Root Step - Sd = (3 - Sn * Sm) / 2 - VRSQRTS.F32 Sd, Sn, Sm
VSHL
Vector Shift Left - Logical/arithmetic shift left - VSHL.S32 Dd, Dm, #imm
VSQRT
Vector Square Root - Sd = sqrtSm - VSQRT.F32 Sd, Sm
VST1
Vector Store One Structure - Store vector register to memory - VST1.32 {Sd}, [Rn]
VST2
Vector Store Two Structures - Store and interleave two values - VST2.32 {Sd, Se}, [Rn]
VST3
Vector Store Three Structures - Store and interleave three values - VST3.32 {Sd, Se, Sf}, [Rn]
VST4
Vector Store Four Structures - Store and interleave four values - VST4.32 {Sd, Se, Sf, Sg}, [Rn]
VSUB
Vector Subtract - Sd = Sn - Sm - VSUB.F32 Sd, Sn, Sm

More in Architecture

  • ARM Assembly LabA working ARM Cortex-M0 assembler and CPU emulator in the browser: real Thumb encodings, N/Z/C/V flags, memory-mapped LEDs/UART/timer, NVIC exception stacking, and graded assembly challenges with automated checking.
  • BootLab 0 → 100Design a reliable bootloader and firmware-update system from reset handoff and flash geometry through image authentication, A/B slots, atomic metadata, trial boot, rollback, recovery, and exhaustive power-cut injection.
  • NVICNVIC for Cortex-M: priority levels, tail-chaining, late-arrival, and vector table. Interactive lab for interrupt latency and preemption.
  • How to Read Any CPUThe six questions that pin down any processor's contract with software: registers, PC and stack, status flags, load/store or memory operands, exception entry, and the memory model. Asked once in the abstract, then answered by every architecture in this section.
  • 8051 / MCS-518051/MCS-51 simulator: 4 register banks, SFR map, bit-addressable RAM, and MOV/CJNE/DJNZ instructions. Observe accumulator and PSW flags.
  • Cortex-AArm Cortex-A model: MMU page tables, exception levels EL0-EL3, and GIC. Interactive simulator for virtual memory translation.
  • AVRAVR architecture: Harvard 8-bit RISC, 32 general-purpose registers, SRAM, and EEPROM. Interactive register and memory access lab.
  • Assembly Across ISAsAssembly for Arm Thumb, RISC-V, AVR, and 8051: mnemonics, addressing modes, and stack frames. Interactive assembler/disassembler.
  • Cortex-RArm Cortex-R real-time core: MPU regions, tightly-coupled memory, and dual-core lockstep. Interactive MPU configuration lab.
  • Architecture MapInteractive architecture map comparing Arm Cortex-M/A/R, RISC-V, AVR, and 8051 ISA, memory maps, and interrupt models.
  • RISC-VRISC-V ISA: RV32I/RV64I base integer set, M/S/U privilege levels, and CSR registers. Interactive instruction execution simulator.
  • ArchitectureArchitecture from reset to reliable products: bootloaders and firmware recovery, Arm Cortex-M/A/R, NVIC, RISC-V, AVR, 8051, and a live Cortex-M0 lab.