Cortex-M
Arm Cortex-M programmers model, Thumb-2 ISA, memory-mapped peripherals, and NVIC interrupt handling. Interactive register and stack simulator.
Arm Cortex-M: The One You Will Actually Ship
Cortex-M is the default architecture for embedded work, and if a reader learns exactly one processor properly this should be it. Its design goal was not peak throughput but deterministic, low-latency response with a programming model that ordinary C can reach - which is why its most distinctive features are all about exceptions rather than arithmetic. The hardware stacks registers for you on interrupt entry, which means a plain C function can be an interrupt handler with no assembly wrapper at all. Coming to it after the 8051, the six programmer's-model questions have familiar shapes and richer answers, and the differences are the lesson: a single flat address space instead of four, load/store instead of memory operands, two stack pointers instead of one, and an interrupt controller that does most of the work a 8051 handler had to do by hand.
How it is built
- Sixteen registers, of which R0 to R12 are general purpose and the last three are structural: R13 is the stack pointer, R14 the link register holding the return address, R15 the program counter. The link register is the significant one - a function call does not push a return address to the stack, it puts it in LR, so a leaf function that calls nothing needs no stack traffic at all to return.
- There are two stack pointers, MSP and PSP, and R13 is whichever one is currently selected. Handler mode always uses MSP; Thread mode can use either. That split is what lets an RTOS give each task its own stack while interrupts run on a single known one, and it is why a stack overflow in a task does not necessarily corrupt the kernel.
- The instruction set is Thumb-2: a mix of 16-bit and 32-bit encodings in one stream, chosen per instruction for density. There is no separate 32-bit mode to switch into - Cortex-M executes only Thumb, which is why the vector table's addresses have their low bit set and why clearing it faults immediately.
- It is a load/store architecture. Data moves between registers and memory only through LDR and STR and their variants; everything else is register to register. The APSR carries N, Z, C and V, and whether an instruction updates them is part of the encoding - ADD leaves flags alone, ADDS sets them, and forgetting the S is a common source of branches that never fire.
- Exception entry is the architecture's centrepiece. On an interrupt the hardware pushes eight registers - R0-R3, R12, LR, PC and xPSR - before the handler runs, so the handler may use those freely and the AAPCS calling convention is satisfied. That is what makes a C function directly usable as an ISR, and it is the single biggest practical difference from the 8051.
- The exception mechanism has optimisations that matter for latency. Tail-chaining skips the unstack-then-restack between two back-to-back interrupts; late arrival lets a higher-priority interrupt take over an entry sequence already in progress. Both reduce worst-case latency, and both are architectural guarantees rather than implementation luck.
- Memory is one flat 4 GB space with an architecturally defined layout: code low, SRAM next, peripherals above that, and the System Control Space near the top holding the NVIC and SCB. There is no MMU. An optional MPU can enforce region permissions but does not translate addresses, so a pointer means the same thing everywhere and there is no per-process address space.
- The vector table is the boot contract. Word zero is the initial stack pointer value, word one is the reset handler's address, and the rest are exception handlers. The processor loads SP from the table before executing a single instruction, which is why a corrupt first word produces a fault before main can possibly run.
Design procedure
- Answer the six questions again on Cortex-M and diff them against your 8051 answers. The differences are the content of this topic.
- Read the vector table in a real startup file and identify the first two words. Understanding that the hardware loads SP from word zero explains most early boot failures.
- Write an ISR as a plain C function and check the disassembly for the absence of a save-and-restore preamble. That absence is the automatic stacking doing its job.
- Find an ADDS in compiled output and work out why the compiler chose the flag-setting form. Flags-as-encoding is the detail that separates reading Thumb from guessing at it.
- Look at where MSP and PSP are selected in an RTOS port. The two-stack split is invisible in bare-metal code and central to everything above it.
- Then go to the NVIC topic. Registers and instructions are half of Cortex-M; the exception controller is the half that decides whether firmware meets its deadlines.
Key terms
- Thumb-2
- Mixed 16- and 32-bit encodings in one stream. Cortex-M executes only Thumb, which is why vector addresses have bit 0 set.
- LR (R14)
- The link register holding the return address. A leaf function returns with no stack traffic at all.
- MSP / PSP
- Two stack pointers. Handler mode always uses MSP; an RTOS gives tasks their own PSP.
- Automatic stacking
- The hardware pushes R0-R3, R12, LR, PC and xPSR on entry, so a C function can be an ISR unmodified.
- EXC_RETURN
- The magic LR value on exception entry. Returning to it, rather than to a normal address, is what unstacks.
- Tail-chaining
- Back-to-back interrupts skip the unstack and restack between them. An architectural latency guarantee.
- S suffix
- Whether an instruction updates flags is part of its encoding. ADD does not; ADDS does.
- Vector table
- Word 0 is the initial SP, word 1 the reset handler. The processor reads SP before executing anything.
Worked example
@ Cortex-M against the 8051, on the same job: toggle a pin in an ISR.
.section .isr_vector
.word _estack @ word 0: hardware loads SP from here
.word Reset_Handler + 1 @ word 1: +1 because Thumb. Clearing
@ that bit faults on the first fetch.
@ A plain C function works as an ISR. No PUSH wall, because the
@ hardware already stacked R0-R3, R12, LR, PC and xPSR for us.
TIM2_IRQHandler:
ldr r0, =GPIOA_ODR @ load/store: memory only via LDR/STR
ldr r1, [r0]
eor r1, r1, #1
str r1, [r0]
bx lr @ lr holds EXC_RETURN, so this unstacks
# What the hardware stacks, and why those eight:
#
# R0 R1 R2 R3 R12 the AAPCS caller-saved set
# LR so the handler can call other functions
# PC where to resume
# xPSR flags and the interrupted state
#
# Exactly the registers a C function is allowed to clobber. That is
# not a coincidence - it is why no assembly wrapper is needed.
# The 8051 comparison, kept from the previous topic:
#
# 8051 Cortex-M
# hardware saves PC only 8 registers
# handler prologue PUSH ACC, PSW... none
# address spaces four one flat
# memory operands yes no - load/store
# stack pointers one two (MSP, PSP)Common pitfalls
General Data Processing & I/O Control
Basic arithmetic, logic, and control instructions available on all Cortex-M cores
- ADC
- Add with Carry - Rd = Rn + Rm + Carry - ADC{S}{cond} Rd, Rn, Rm
- ADD
- Add - Rd = Rn + operand2 - ADD{S}{cond} Rd, Rn, #imm | Rm
- AND
- Bitwise AND - Rd = Rn & Rm - AND{S}{cond} Rd, Rn, Rm
- ASR
- Arithmetic Shift Right - Rd = Rn >> imm (sign-extended) - ASR{S}{cond} Rd, Rn, #imm | Rs
- B
- Branch - PC = label - B{cond} label
- BIC
- Bit Clear - Rd = Rn & ~Rm - BIC{S}{cond} Rd, Rn, Rm
- BKPT
- Breakpoint - Enter debug state - BKPT #imm
- BL
- Branch with Link - LR = PC+4; PC = label - BL label
- BLX
- Branch with Link and Exchange - LR = PC+4; PC = Rm; T bit = Rm[0] - BLX Rm | label
- BX
- Branch and Exchange - PC = Rm; T bit = Rm[0] - BX Rm
- CBNZ
- Compare and Branch on Non-Zero - if (Rn != 0) PC = label - CBNZ Rn, label
- CBZ
- Compare and Branch on Zero - if (Rn == 0) PC = label - CBZ Rn, label
- CLZ
- Count Leading Zeros - Rd = number of leading zeros in Rn - CLZ Rd, Rn
- CMN
- Compare Negative - Flags = Rn + Rm - CMN{cond} Rn, Rm
- CMP
- Compare - Flags = Rn - operand2 - CMP{cond} Rn, Rm | #imm
- CPS
- Change Processor State - Set/clear I or F interrupt masks - CPSID i | f | CPSIE i | f
- DMB
- Data Memory Barrier - Wait for all memory accesses to complete - DMB {option}
- DSB
- Data Synchronization Barrier - Complete all memory accesses - DSB {option}
- EOR
- Exclusive OR - Rd = Rn ^ Rm - EOR{S}{cond} Rd, Rn, Rm
- ISB
- Instruction Synchronization Barrier - Flush pipeline and refetch instructions - ISB {option}
- LDMIA
- Load Multiple Increment After - Load registers from sequential memory - LDMIA Rn!, {Rlist}
- LDR
- Load Register - Rd = Memory[Rn + offset] - LDR{cond} Rd, [Rn, #offset]
- LDRB
- Load Register Byte - Rd = zero_extend(Memory[Rn + offset]) - LDRB{cond} Rd, [Rn, #offset]
- LDRH
- Load Register Halfword - Rd = zero_extend(Memory[Rn + offset]) - LDRH{cond} Rd, [Rn, #offset]
- LDRSB
- Load Register Signed Byte - Rd = sign_extend(Memory[Rn + offset]) - LDRSB{cond} Rd, [Rn, #offset]
- LDRSH
- Load Register Signed Halfword - Rd = sign_extend(Memory[Rn + offset]) - LDRSH{cond} Rd, [Rn, #offset]
- LSL
- Logical Shift Left - Rd = Rn << imm - LSL{S}{cond} Rd, Rn, #imm | Rs
- LSR
- Logical Shift Right - Rd = Rn >> imm (zero-filled) - LSR{S}{cond} Rd, Rn, #imm | Rs
- MOV
- Move - Rd = operand2 - MOV{S}{cond} Rd, #imm | Rm
- MUL
- Multiply - Rd = Rn * Rm (lower 32 bits) - MUL{S}{cond} Rd, Rn, Rm
- MVN
- Move Not - Rd = ~Rm - MVN{S}{cond} Rd, Rm
- NOP
- No Operation - None - NOP
- ORR
- Bitwise OR - Rd = Rn | Rm - ORR{S}{cond} Rd, Rn, Rm
- POP
- Pop Registers from Stack - Load registers from [SP]; SP += 4*n - POP {Rlist}
- PUSH
- Push Registers to Stack - SP -= 4*n; Store registers to [SP] - PUSH {Rlist}
- REV
- Reverse Byte Order - Rd = byte_reverse(Rn) - REV Rd, Rn
- REV16
- Reverse Byte Order in Halfwords - Reverse bytes in each 16-bit half - REV16 Rd, Rn
- REVSH
- Reverse Signed Halfword - Reverse bytes in low halfword, sign-extend - REVSH Rd, Rn
- ROR
- Rotate Right - Rd = Rn rotated right by Rm - ROR{S}{cond} Rd, Rn, Rm
- RSB
- Reverse Subtract - Rd = Rm - Rn - RSB{S}{cond} Rd, Rn, Rm
- SDIV
- Signed Divide - Rd = Rn / Rm (signed) - SDIV{cond} Rd, Rn, Rm
- STMIA
- Store Multiple Increment After - Store registers to sequential memory - STMIA Rn!, {Rlist}
- STR
- Store Register - Memory[Rn + offset] = Rd - STR{cond} Rd, [Rn, #offset]
- STRB
- Store Register Byte - Memory[Rn + offset] = Rd[7:0] - STRB{cond} Rd, [Rn, #offset]
- STRH
- Store Register Halfword - Memory[Rn + offset] = Rd[15:0] - STRH{cond} Rd, [Rn, #offset]
- SUB
- Subtract - Rd = Rn - operand2 - SUB{S}{cond} Rd, Rn, #imm | Rm
- SVC
- Supervisor Call - Trigger SVC exception - SVC #imm
- SXTB
- Sign-Extend Byte - Rd = sign_extend(Rn[7:0]) - SXTB Rd, Rn
- SXTH
- Sign-Extend Halfword - Rd = sign_extend(Rn[15:0]) - SXTH Rd, Rn
- TST
- Test - Flags = Rn & Rm - TST{cond} Rn, Rm
- UDIV
- Unsigned Divide - Rd = Rn / Rm (unsigned) - UDIV{cond} Rd, Rn, Rm
- UXTB
- Zero-Extend Byte - Rd = zero_extend(Rn[7:0]) - UXTB Rd, Rn
- UXTH
- Zero-Extend Halfword - Rd = zero_extend(Rn[15:0]) - UXTH Rd, Rn
- WFE
- Wait For Event - Wait for event signal - WFE
- WFI
- Wait For Interrupt - Wait for interrupt - WFI
- YIELD
- Yield - Hint to processor for thread switching - YIELD
Advanced Data Processing & Bit Field
Enhanced instructions for complex data manipulation and bit field operations
- ADR
- Address to Register (PC-relative) - Rd = PC + offset to label - ADR Rd, label
- BFC
- Bit Field Clear - Clear width bits starting at lsb - BFC Rd, #lsb, #width
- BFI
- Bit Field Insert - Copy width bits from Rn to Rd at lsb - BFI Rd, Rn, #lsb, #width
- CDP
- Coprocessor Data Processing - Coprocessor-specific operation - CDP p#, opcode, CRd, CRn, CRm
- DBG
- Debug - Debug hint - DBG #imm
- LDC
- Load Coprocessor - Load coprocessor register from memory - LDC p#, CRd, [Rn, #offset]
- LDM
- Load Multiple - Load multiple registers - LDM{mode} Rn, {Rlist}
- LDRD
- Load Register Dual - Load 64-bit value to Rd:Rt - LDRD Rd, Rt, [Rn, #offset]
- LDMDB
- Load Multiple Decrement Before - Decrement then load - LDMDB Rn!, {Rlist}
- MCR
- Move to Coprocessor from Register - Copy Rd to coprocessor register - MCR p#, opcode, Rd, CRn, CRm
- MCRR
- Move to Coprocessor from Two Registers - Copy Rd:Rt to coprocessor - MCRR p#, opcode, Rd, Rt, CRm
- MLA
- Multiply Accumulate - Rd = (Rn * Rm) + Ra - MLA{S}{cond} Rd, Rn, Rm, Ra
- MLS
- Multiply Subtract - Rd = Ra - (Rn * Rm) - MLS{cond} Rd, Rn, Rm, Ra
- MRC
- Move to Register from Coprocessor - Copy coprocessor register to Rd - MRC p#, opcode, Rd, CRn, CRm
- MRRC
- Move to Two Registers from Coprocessor - Copy coprocessor to Rd:Rt - MRRC p#, opcode, Rd, Rt, CRm
- MRS
- Move to Register from Special Register - Rd = special register - MRS Rd, spec_reg
- MSR
- Move to Special Register from Register - special register = Rn - MSR spec_reg, Rn
- PLD
- Prefetch Data - Hint to prefetch memory - PLD [Rn, #offset]
- PLI
- Prefetch Instruction - Hint to prefetch instruction - PLI [Rn, #offset]
- RBIT
- Reverse Bits - Rd = bit_reverse(Rn) - RBIT Rd, Rn
- SBFX
- Signed Bit Field Extract - Extract width bits from lsb, sign-extend - SBFX Rd, Rn, #lsb, #width
- SDIV
- Signed Divide - Rd = Rn / Rm (signed) - SDIV{cond} Rd, Rn, Rm
- SEL
- Select Bytes - Select bytes from Rn or Rm based on GE - SEL Rd, Rn, Rm
- SEV
- Send Event - Generate event signal - SEV
- SMC
- Secure Monitor Call - Enter secure monitor mode - SMC #imm
- SMMLA
- Signed Most Significant Word Multiply Accumulate - Rd = Ra + (Rn * Rm)[63:32] - SMMLA{R} Rd, Rn, Rm, Ra
- SMMLS
- Signed Most Significant Word Multiply Subtract - Rd = Ra - (Rn * Rm)[63:32] - SMMLS{R} Rd, Rn, Rm, Ra
- SMMUL
- Signed Most Significant Word Multiply - Rd = (Rn * Rm)[63:32] - SMMUL{R} Rd, Rn, Rm
- STM
- Store Multiple - Store multiple registers - STM{mode} Rn, {Rlist}
- STRD
- Store Register Dual - Store 64-bit value from Rd:Rt - STRD Rd, Rt, [Rn, #offset]
- STREX
- Store Exclusive - Attempt exclusive store, Rd = success/fail - STREX Rd, Rn, [Rm]
- STREXB
- Store Exclusive Byte - Exclusive byte store - STREXB Rd, Rn, [Rm]
- STREXH
- Store Exclusive Halfword - Exclusive halfword store - STREXH Rd, Rn, [Rm]
- STRT
- Store Register Unprivileged - Store as unprivileged access - STRT Rd, [Rn]
- SUBS
- Subtract with flags - Rd = Rn - Rm; update flags - SUBS Rd, Rn, Rm
- SXTAB
- Sign-Extend and Add Byte - Rd = Rn + sign_extend(Rm[7:0]) - SXTAB Rd, Rn, Rm
- SXTAB16
- Sign-Extend and Add Two Halfwords - Sign-extend bytes to halfwords and add - SXTAB16 Rd, Rn, Rm
- SXTAH
- Sign-Extend and Add Halfword - Rd = Rn + sign_extend(Rm[15:0]) - SXTAH Rd, Rn, Rm
- UBFX
- Unsigned Bit Field Extract - Extract width bits from lsb, zero-extend - UBFX Rd, Rn, #lsb, #width
- UDIV
- Unsigned Divide - Rd = Rn / Rm (unsigned) - UDIV{cond} Rd, Rn, Rm
- UXTAB
- Zero-Extend and Add Byte - Rd = Rn + zero_extend(Rm[7:0]) - UXTAB Rd, Rn, Rm
- UXTAB16
- Zero-Extend and Add Two Halfwords - Zero-extend bytes to halfwords and add - UXTAB16 Rd, Rn, Rm
- UXTAH
- Zero-Extend and Add Halfword - Rd = Rn + zero_extend(Rm[15:0]) - UXTAH Rd, Rn, Rm
DSP Instructions (SIMD & Fast MAC)
Digital Signal Processing instructions with SIMD and Multiply-Accumulate operations
- QADD
- Saturating Add - Rd = saturate(Rn + Rm) - QADD{cond} Rd, Rn, Rm
- QDADD
- Saturating Double and Add - Rd = Rn + saturate(Rm * 2) - QDADD{cond} Rd, Rn, Rm
- QDSUB
- Saturating Double and Subtract - Rd = Rn - saturate(Rm * 2) - QDSUB{cond} Rd, Rn, Rm
- QSUB
- Saturating Subtract - Rd = saturate(Rn - Rm) - QSUB{cond} Rd, Rn, Rm
- SADD16
- Signed Add 16 - Rd[15:0] = Rn[15:0] + Rm[15:0]; Rd[31:16] = Rn[31:16] + Rm[31:16] - SADD16 Rd, Rn, Rm
- SADD8
- Signed Add 8 - Four parallel 8-bit signed additions - SADD8 Rd, Rn, Rm
- SASX
- Signed Add and Subtract with Exchange - Rd[31:16] = Rn[31:16] + Rm[15:0]; Rd[15:0] = Rn[15:0] - Rm[31:16] - SASX Rd, Rn, Rm
- SHADD16
- Signed Halving Add 16 - Rd[15:0] = (Rn[15:0] + Rm[15:0]) >> 1 - SHADD16 Rd, Rn, Rm
- SHASX
- Signed Halving Add and Subtract with Exchange - Halving add/subtract with exchange - SHASX Rd, Rn, Rm
- SHSAX
- Signed Halving Subtract and Add with Exchange - Halving subtract/add with exchange - SHSAX Rd, Rn, Rm
- SHSUB16
- Signed Halving Subtract 16 - Rd[15:0] = (Rn[15:0] - Rm[15:0]) >> 1 - SHSUB16 Rd, Rn, Rm
- SHSUB8
- Signed Halving Subtract 8 - Four parallel halving 8-bit subtractions - SHSUB8 Rd, Rn, Rm
- SMLA
- Signed Multiply Accumulate - Rd = Ra + (Rn * Rm) with saturation - SMLA{B|T}{cond} Rd, Rn, Rm, Ra
- SMLAD
- Signed Multiply Accumulate Dual - Rd = Ra + Rn[15:0]*Rm[15:0] + Rn[31:16]*Rm[31:16] - SMLAD{X}{cond} Rd, Rn, Rm, Ra
- SMLAL
- Signed Multiply Accumulate Long - 64-bit accumulate: RdHi:RdLo += Rn * Rm - SMLAL{cond} RdLo, RdHi, Rn, Rm
- SMLALD
- Signed Multiply Accumulate Long Dual - 64-bit dual multiply accumulate - SMLALD{X}{cond} RdLo, RdHi, Rn, Rm
- SMLSD
- Signed Multiply Subtract Dual - Rd = Ra + Rn[15:0]*Rm[15:0] - Rn[31:16]*Rm[31:16] - SMLSD{X}{cond} Rd, Rn, Rm, Ra
- SMLSLD
- Signed Multiply Subtract Long Dual - 64-bit dual multiply subtract - SMLSLD{X}{cond} RdLo, RdHi, Rn, Rm
- SMMLA
- Signed Most Significant Word Multiply Accumulate - Rd = Ra + (Rn * Rm)[63:32] - SMMLA{R}{cond} Rd, Rn, Rm, Ra
- SMUAD
- Signed Multiply Dual - Rd = Rn[15:0]*Rm[15:0] + Rn[31:16]*Rm[31:16] - SMUAD{X}{cond} Rd, Rn, Rm
- SMUSD
- Signed Multiply Subtract Dual - Rd = Rn[15:0]*Rm[15:0] - Rn[31:16]*Rm[31:16] - SMUSD{X}{cond} Rd, Rn, Rm
- SSAT
- Signed Saturate - Rd = saturate_signed(Rn, imm) - SSAT Rd, #imm, Rn
- SSAT16
- Signed Saturate 16 - Saturate both halfwords to signed imm bits - SSAT16 Rd, #imm, Rn
- SSAX
- Signed Subtract and Add with Exchange - Rd[31:16] = Rn[31:16] - Rm[15:0]; Rd[15:0] = Rn[15:0] + Rm[31:16] - SSAX Rd, Rn, Rm
- SSUB16
- Signed Subtract 16 - Two parallel 16-bit signed subtractions - SSUB16 Rd, Rn, Rm
- SSUB8
- Signed Subtract 8 - Four parallel 8-bit signed subtractions - SSUB8 Rd, Rn, Rm
- UADD16
- Unsigned Add 16 - Two parallel 16-bit unsigned additions - UADD16 Rd, Rn, Rm
- UADD8
- Unsigned Add 8 - Four parallel 8-bit unsigned additions - UADD8 Rd, Rn, Rm
- UASX
- Unsigned Add and Subtract with Exchange - Rd[31:16] = Rn[31:16] + Rm[15:0]; Rd[15:0] = Rn[15:0] - Rm[31:16] - UASX Rd, Rn, Rm
- UHADD16
- Unsigned Halving Add 16 - Rd[15:0] = (Rn[15:0] + Rm[15:0]) >> 1 - UHADD16 Rd, Rn, Rm
- UHASX
- Unsigned Halving Add and Subtract with Exchange - Halving add/sub with exchange - UHASX Rd, Rn, Rm
- UHSAX
- Unsigned Halving Subtract and Add with Exchange - Halving subtract/add with exchange - UHSAX Rd, Rn, Rm
- UHSUB16
- Unsigned Halving Subtract 16 - Rd[15:0] = (Rn[15:0] - Rm[15:0]) >> 1 - UHSUB16 Rd, Rn, Rm
- UHSUB8
- Unsigned Halving Subtract 8 - Four parallel halving 8-bit subtractions - UHSUB8 Rd, Rn, Rm
- UMAAL
- Unsigned Multiply Accumulate Accumulate Long - RdHi:RdLo = Rn + Rm + (Rn * Rm) - UMAAL RdLo, RdHi, Rn, Rm
- UMLAL
- Unsigned Multiply Accumulate Long - RdHi:RdLo += Rn * Rm (unsigned) - UMLAL{cond} RdLo, RdHi, Rn, Rm
- UMULL
- Unsigned Multiply Long - RdHi:RdLo = Rn * Rm (unsigned 64-bit) - UMULL{cond} RdLo, RdHi, Rn, Rm
- UQADD16
- Unsigned Saturating Add 16 - Two parallel saturating 16-bit additions - UQADD16 Rd, Rn, Rm
- UQADD8
- Unsigned Saturating Add 8 - Four parallel saturating 8-bit additions - UQADD8 Rd, Rn, Rm
- UQASX
- Unsigned Saturating Add and Subtract with Exchange - Saturating add/sub with exchange - UQASX Rd, Rn, Rm
- UQSAX
- Unsigned Saturating Subtract and Add with Exchange - Saturating subtract/add with exchange - UQSAX Rd, Rn, Rm
- UQSUB16
- Unsigned Saturating Subtract 16 - Two parallel saturating 16-bit subtractions - UQSUB16 Rd, Rn, Rm
- UQSUB8
- Unsigned Saturating Subtract 8 - Four parallel saturating 8-bit subtractions - UQSUB8 Rd, Rn, Rm
- USAD8
- Unsigned Sum of Absolute Differences 8 - Rd = sum |Rn[i] - Rm[i]| for i=0..3 - USAD8 Rd, Rn, Rm
- USADA8
- Unsigned Sum of Absolute Differences 8 Accumulate - Rd = Ra + sum |Rn[i] - Rm[i]| - USADA8 Rd, Rn, Rm, Ra
- USAT
- Unsigned Saturate - Rd = saturate_unsigned(Rn, imm) - USAT Rd, #imm, Rn
- USAT16
- Unsigned Saturate 16 - Saturate both halfwords to unsigned imm bits - USAT16 Rd, #imm, Rn
- USAX
- Unsigned Subtract and Add with Exchange - Rd[31:16] = Rn[31:16] - Rm[15:0]; Rd[15:0] = Rn[15:0] + Rm[31:16] - USAX Rd, Rn, Rm
- USUB16
- Unsigned Subtract 16 - Two parallel 16-bit unsigned subtractions - USUB16 Rd, Rn, Rm
- USUB8
- Unsigned Subtract 8 - Four parallel 8-bit unsigned subtractions - USUB8 Rd, Rn, Rm
- UXTB16
- Unsigned Extend Byte 16 - Zero-extend Rn[7:0] and Rn[23:16] to halfwords - UXTB16 Rd, Rn
Floating Point Instructions
Single-precision floating point operations (VFPv4-SP)
- VABS
- Floating-Point Absolute - Sd = |Sm| - VABS.F32 Sd, Sm
- VADD
- Floating-Point Add - Sd = Sn + Sm - VADD.F32 Sd, Sn, Sm
- VCMP
- Floating-Point Compare - Compare Sn with Sm - VCMP.F32 Sn, Sm
- VCMP
- Floating-Point Compare with Zero - Compare Sn with 0.0 - VCMP.F32 Sn, #0
- VCVT
- Floating-Point Convert - Convert float to int or int to float - VCVT.S32.F32 Sd, Sm | VCVT.F32.S32 Sd, Sm
- VCVTR
- Floating-Point Convert with Rounding - Convert with current rounding mode - VCVTR.S32.F32 Sd, Sm
- VDIV
- Floating-Point Divide - Sd = Sn / Sm - VDIV.F32 Sd, Sn, Sm
- VLD1
- Vector Load One Structure - Load from memory to vector register - VLD1.32 {Sd}, [Rn]
- VLD2
- Vector Load Two Structures - Load and deinterleave two values - VLD2.32 {Sd, Se}, [Rn]
- VLD3
- Vector Load Three Structures - Load and deinterleave three values - VLD3.32 {Sd, Se, Sf}, [Rn]
- VLD4
- Vector Load Four Structures - Load and deinterleave four values - VLD4.32 {Sd, Se, Sf, Sg}, [Rn]
- VMOV
- Vector Move - Move value to vector register - VMOV.F32 Sd, Sm | VMOV Sd, #imm
- VMRS
- Vector Move to Register from Special - Copy FPSCR to general register - VMRS Rd, FPSCR
- VMSR
- Vector Move to Special from Register - Copy general register to FPSCR - VMSR FPSCR, Rd
- VMUL
- Vector Multiply - Sd = Sn * Sm - VMUL.F32 Sd, Sn, Sm
- VNEG
- Vector Negate - Sd = -Sm - VNEG.F32 Sd, Sm
- VNMLA
- Vector Negative Multiply Accumulate - Sd = -(Sd + Sn * Sm) - VNMLA.F32 Sd, Sn, Sm
- VNMLS
- Vector Negative Multiply Subtract - Sd = -(Sd - Sn * Sm) - VNMLS.F32 Sd, Sn, Sm
- VNMUL
- Vector Negative Multiply - Sd = -(Sn * Sm) - VNMUL.F32 Sd, Sn, Sm
- VPOP
- Vector Pop - Pop from stack to vector registers - VPOP {Slist}
- VPUSH
- Vector Push - Push vector registers to stack - VPUSH {Slist}
- VRADDHN
- Vector Rounding Add Narrow - Add and narrow with rounding - VRADDHN.D16 Dd, Dn, Dm
- VRECPE
- Vector Reciprocal Estimate - Sd ~= 1/Sm (approximate) - VRECPE.F32 Sd, Sm
- VRECPS
- Vector Reciprocal Step - Sd = 2 - Sn * Sm - VRECPS.F32 Sd, Sn, Sm
- VRINTA
- Vector Round to Integral (Away) - Round away from zero - VRINTA.F32 Sd, Sm
- VRINTM
- Vector Round to Integral (Minus) - Round toward -infinity - VRINTM.F32 Sd, Sm
- VRINTN
- Vector Round to Integral (Nearest) - Round to nearest (ties to even) - VRINTN.F32 Sd, Sm
- VRINTP
- Vector Round to Integral (Plus) - Round toward +infinity - VRINTP.F32 Sd, Sm
- VRINTR
- Vector Round to Integral (with Rounding) - Round using FPSCR rounding mode - VRINTR.F32 Sd, Sm
- VRINTX
- Vector Round to Integral (Exact) - Round to integral, set IXF if inexact - VRINTX.F32 Sd, Sm
- VRINTZ
- Vector Round to Integral (Zero) - Round toward zero (truncate) - VRINTZ.F32 Sd, Sm
- VRSHR
- Vector Round Shift Right - Shift right with rounding - VRSHR.S32 Dd, Dm, #imm
- VRSQRT
- Vector Reciprocal Square Root Estimate - Sd ~= 1/sqrtSm - VRSQRT.F32 Sd, Sm
- VRSQRTS
- Vector Reciprocal Square Root Step - Sd = (3 - Sn * Sm) / 2 - VRSQRTS.F32 Sd, Sn, Sm
- VSHL
- Vector Shift Left - Logical/arithmetic shift left - VSHL.S32 Dd, Dm, #imm
- VSQRT
- Vector Square Root - Sd = sqrtSm - VSQRT.F32 Sd, Sm
- VST1
- Vector Store One Structure - Store vector register to memory - VST1.32 {Sd}, [Rn]
- VST2
- Vector Store Two Structures - Store and interleave two values - VST2.32 {Sd, Se}, [Rn]
- VST3
- Vector Store Three Structures - Store and interleave three values - VST3.32 {Sd, Se, Sf}, [Rn]
- VST4
- Vector Store Four Structures - Store and interleave four values - VST4.32 {Sd, Se, Sf, Sg}, [Rn]
- VSUB
- Vector Subtract - Sd = Sn - Sm - VSUB.F32 Sd, Sn, Sm
More in Architecture
- ARM Assembly LabA working ARM Cortex-M0 assembler and CPU emulator in the browser: real Thumb encodings, N/Z/C/V flags, memory-mapped LEDs/UART/timer, NVIC exception stacking, and graded assembly challenges with automated checking.
- BootLab 0 → 100Design a reliable bootloader and firmware-update system from reset handoff and flash geometry through image authentication, A/B slots, atomic metadata, trial boot, rollback, recovery, and exhaustive power-cut injection.
- NVICNVIC for Cortex-M: priority levels, tail-chaining, late-arrival, and vector table. Interactive lab for interrupt latency and preemption.
- How to Read Any CPUThe six questions that pin down any processor's contract with software: registers, PC and stack, status flags, load/store or memory operands, exception entry, and the memory model. Asked once in the abstract, then answered by every architecture in this section.
- 8051 / MCS-518051/MCS-51 simulator: 4 register banks, SFR map, bit-addressable RAM, and MOV/CJNE/DJNZ instructions. Observe accumulator and PSW flags.
- Cortex-AArm Cortex-A model: MMU page tables, exception levels EL0-EL3, and GIC. Interactive simulator for virtual memory translation.
- AVRAVR architecture: Harvard 8-bit RISC, 32 general-purpose registers, SRAM, and EEPROM. Interactive register and memory access lab.
- Assembly Across ISAsAssembly for Arm Thumb, RISC-V, AVR, and 8051: mnemonics, addressing modes, and stack frames. Interactive assembler/disassembler.
- Cortex-RArm Cortex-R real-time core: MPU regions, tightly-coupled memory, and dual-core lockstep. Interactive MPU configuration lab.
- Architecture MapInteractive architecture map comparing Arm Cortex-M/A/R, RISC-V, AVR, and 8051 ISA, memory maps, and interrupt models.
- RISC-VRISC-V ISA: RV32I/RV64I base integer set, M/S/U privilege levels, and CSR registers. Interactive instruction execution simulator.
- ArchitectureArchitecture from reset to reliable products: bootloaders and firmware recovery, Arm Cortex-M/A/R, NVIC, RISC-V, AVR, 8051, and a live Cortex-M0 lab.