Have you ever wondered when you are woking on an Arduino project how does everything works, why sometimes it takes so long to flash or compile the code..., Or when you work an ST project on the CUBE-IDE and you hit run or Debug what is happening on the backend, Or you are just fed up with using a bloated IDE
well, let's figure out to program an STM32 and set up its environment from scratch.
I am using a STM32F4RET6-NUCLEO board, it has a CORTEX-M4 micro-processor and 00 FLASH and 00 SRAM.
This article will explore how to set up a development environment from scratch, starting with a bare main.c file as our main application (a blinky sketch).
A development should contains these elements :
code editor: a piece of software to write your code.toolchain: tools needed to compile, assemble ...
-Startup file: describes the Request handler that the MCU calls when an IT happens, describes the VT and implements theReset_handler.linker script: describes you memory layout and how to set up different sections
In this article i will focus how to setup an Development environment from scratch without the need of an IDE.
I will assume that you are using ubuntu as your operating system.
- Article 1: Development Environment from Scratch for an STM32
- An IDE VS Code Editor
- Embedded System Code Structure
- Tool-Chains
- Conclusion
#todo
if you worked on the STM32-Cube IDE, you would find this structure :
Src dir: contains all the source files.c(main.c)Includes dir: contains all the headers.hDebug dir: contains the build outputStartup dir: contains the start up file.ldfile extension: linker script
#todo: insert image
There are other directories like driver and core, i am not going to focus on them for the moment.
i have already created the simple main.c file: a blinky application ( we will explore it later ).
Now How to create a linker script and a startup file for an embedded target.
The linker script is a piece of code that describes the memory layout of your project
It describes the physical memory size (Flash and SRAM), also it describes the sections present in your code.
Flash: this memory contains the binary code you uploaded and it's going to run (volatile).SRAM: this is the ram that would hold your data at runtime (non-volatile)
- Let's first start with flash:
As said before the flash holds the application binary code, it is divided into sections:
- `.text` section : contains the application code.
- `.data` section : contains the (global/static) initialized data
- `.rodata` section : contains the constant data.
- `.bss` section : contains the (global/static) uninitialized data.
- The
SRAM(static memory):
The SRAM holds the data at runtime:
- `.data` section : same as the `flash`, the `.data` of the `flash` are copied to `.data` section of `SRAM` at boot-time.
- `.bss` section : same as the `flash`, the `.bss` of the `flash` are copied and initialized to `0` to `.bss` section of `SRAM` at boot-time.
- `stack` section : contains the local data that are statically allocated and function calls at runtime.
- `heap` section : contains dynamic allocated memory regions, at runtime.
Peripherals: a memory area, describes where the MCU peripherals are mapped, it is used to control the peripherals through registers.
In a PC you have in the mother-board, CPU, DISK and MEMORY stick, each one in dedicated area and connected through each other via buses.
In an STM32 (STM32F4) there is a 4Gb of addressable memory space, this memory area is divided into sections, each section has a defined boundary address space (start and end), these sections represents the different memory component of an embedded target.
This space is not a physical memory block but rather a way of organizing how different types of physical memory and peripherals are accessed by the processor.
This information is present in the Memory mapping section in the DATASHEET.
TODO: Insert Picture
This means to access SRAM or FLASH, they are mapped to a specific location in this memory space.
These mappings are not standard and changes from MCU to MCU, for that we need to specify the location of each region, it's size and it's contents (sections).
This is where the linker comes into play.
The linker is responsible for assigning memory addresses to .text and .data sections of a program, ensuring that they are correctly placed in the appropriate memory regions as defined by the microcontroller’s (MCU's) memory map. This process is guided by a linker script.
A linker script is a text file that details how different sections of an object file should be combined to create the final output file. It specifies unique absolute addresses for these sections based on the address information provided within the script.
Key aspects of a linker script include:
- Memory Layout : Describes how memory is organized and assigns addresses to different sections.
- Code and Data Addresses : Defines where code and data should be placed in memory and their sizes.
Linker scripts are written using GNU linker command language and have the .ld file extension.
During the linking phase, you must supply the linker script to the linker using the -T option.
The linker performs additional roles, which will be discussed later.
writing a linker script is very easy.
- This is a basic template how a linker script should appear.
ENTRY (_symbol_name_)
MEMORY {
memory-name (att) :ORIGIN = adr,LENGTH = size
}
SECTIONS {
.section_name:{
*(.section_name)
}>(VMA) AT> (LMA)
}This command sets the entry-point address information in the header .elf file (final executable) to _symbol_name_.
Usually the _symbol_name_ refers to the first function to be executed by the CPU, in ST this function is called Reset_Handler.
- let's replace the
ENTRYfunction with our suitable Handler:
ENTRY(Reset_Handler)This command allows you to describe the different memories present in the target and their start address and size information.
This information also helps the linker to calculate total code and memory consumed so far and throw error if data, code, heap or stack areas cannot fit into the available sizes
memory-name (perm) :ORIGIN = adr,LENGTH = size-
memory-name: defines a name of the memory region which will be later referenced by other parts of thelinker script. -
adr: defines origin address, start of the memory region -
size: defines length of information -
att: defines the attribute list of the memory region (actions performed Read/write access ...). -
Let's write the
MEMORYcommand based on theSTM32F4RET6:
This board has: -
512K of
Flash -
96K of
SRAM
if you have other SRAM regions you can merge them together or declare them separately.
MEMORY{
FLASH (rx) :ORIGIN =0x08000000,LENGTH =512K
SRAM (rwx) :ORIGIN =0x20000000,LENGTH =96K
}RX: in flash means that, its a Read only memory region and contains executable code.W: means its readable and writable. this ensures that code resides in a protected memory region and that data can be modified as needed.
This section is responsible for creating the output sections present in the final executable. by merging all of the input section.
lets say we have .c file, let's compile it using -c to obtain .o file representing the binary format.
gcc -c main.c -o main.o Notice here, we did not command the gcc to link, we just compiled and translated the .c file into binary format.
we will explore this in a later chapter.
Imagine now we have multiple .o files. Each binary file has section like .text, .data.
In the SECTIONS we telling the linker to combine all of the .text section together, etc..
SECTIONS{
/* here we are instructing the linker to merge all the .text in a section called .text*/
.text{
*(.text)
/* onther way is to specify the the files we want to get the .text section from */
main.o file1.o // ...
}
}we can do that for all the other sections.
SECTIONS
{
.text : {
*(.text)
*(.rodata)
}
.data : {
*(.data)
}
.bss : {
*(.bss)
}
}Note i just added the
.rodataunder the.textsections
We can control also the oder the section in the final elf (who shows up first), this is important, will see that later.
- now how to control the placement of a section in a memory region, for example set the
.textsection to be in FLASH memory region ?
we can do that using keywords:
VMA:virtual memory address (static memory), it could be anotherSRAMyou defined (SRAM,SRAM2).
LMA: load memory address (physical memory),FLASHor an external Flash memory
ATcommand.
using these variations, we can instruct the linker where to put the different sections, (FLASH,SRAM,EXternal FLASH) it depends on your application.
- To instruct the linker to put the section in LMA
.section_name{
}> FLASH- To instruct the linker relocate the section in from LMA to VMA
.section_name{
}> SRAM AT> FLASH- To instruct the linker to put the section in VMA
.section_name{
}> SRAMwe obtain this code now:
SECTIONS
{
/* text section */
.text : {
*(.text)
*(.rodata)
}>FLASH
/* data section */
.data : {
*(.data)
}> SRAM AT> FLASH
/* uninitialized data */
.bss : {
*(.bss)
}> SRAM
}in linker script there are other utilities we need in the next section used in the SECTIONS
.(dot) operator:
the . operator is counter that keeps record of used memory.
for example we can measure the end of the .text section
.text{
/* specifies the beginning of the address */
. ;
*(.text)
*(.rodata)
/* store the end of .text section in the symbol end_of_text */
end_of_text = . ;
}- ALIGN(4):
this command ensures that memory region is a aligned by
4bytes.
the memory address values is mutiple of4(0x08000000,0x08000004,0x08000008).
Why aligning an address.
The CPU (CORTEX-M4)32 bit register, meaning the processor access the memory in chunks of4bytes, when data isALIGN(4)the CPU can fetch the data in one memory cycle, increasing efficiency.
lest's takeend_of_textis a variable containing the end text data region address. that address coud be0x08000082(not aligned).
If the next section, say.data, starts immediately at0x08000083, this would create analignment issue. Since0x08000083is not a multiple of4, the CPU would need extra steps to access the data efficiently, leading to performance degradation
.text{
/* specifies the beginning of the address */
. ;
*(.text)
*(.rodata)
/* store the end of .text section in the symbol end_of_text */
. = ALIGN(4) ;
end_of_text = . ;
}arm-none-gcc -nostdlib -T stm32_ls.ld *.o -o prog.elf-nostdlib: specifies that we are not using theglibcthe C standard libary. ( gcc by default assumes you are usinglibc)-Map: creates a file with a detailed overview of the memory layout, including the sizes and locations of all sections and symbols.
the linker script is an essential file to describe how you want to control your memory.
In the next Section we will see what is start up file and how to write one.
You have come up this far, Nice !!!
so the million question, what is a startup file.
A start up file is a piece of code that implements the request handlers (don't worry wil see that later), launches the Reset_Handler and describes the vector table (VT).
One key component to remember is that the startup file is the one that calls the main.
You might find an implementation of the startup file written in Assembly (yeah you read that right ASSEMBLY), but this time we are going to take things easily and write it in C.
this was a general overview about the startup file. Now let's dive in:
When working with embedded targets like the STM32 microcontroller, we heavily rely on interrupts. Interrupts are a crucial mechanism that allows the microcontroller to pause the current execution flow to handle higher-priority tasks (will see in a later section how interrupts works).
These interrupts can take various forms:
- Exceptions (System Level): Such as a division by zero error.
- External Interrupts: Like a button press.
- Non-Maskable Interrupts (NMI): An interrupt that cannot be disabled.
When an interrupt occurs, the microcontroller executes a specific task associated with it, known as a Handler. Each interrupt has a corresponding handler, such as EXTI0_IRQHandler(External interrupt), and these handlers are unique—meaning each interrupt can only have one associated handler.
Interrupts are specific to the microcontroller's architecture, meaning you cannot create your own interrupts or assign interrupts to peripherals arbitrarily. Each interrupt is assigned a priority, which is typically documented in the manufacturer's reference manual.
For example, in the STM32F4 series micro-controllers, the reference manual provides detailed information about the available interrupts and their priorities. This information is presented in what is known as the Vector Table (VT).
The Vector Table is a critical component of the micro-controller's architecture. It is essentially a table of pointers that point to the start addresses of the interrupt handlers. When an interrupt occurs, the microcontroller looks up the corresponding address in the Vector Table and jumps to that address to execute the handler.
Here's an example from the STM32F4RE reference manual, where you can see the Vector Table layout:
(Insert image here)
The Vector Table is located at the beginning of the microcontroller's memory map(FLASH), starting from address 0x00000000.
Usually the first entry in the table is the initial stack pointer, in other MCU's you might find that spot as reserved, followed by the Reset_Handler address, which points to the start of the Reset_Handler. After that, each entry corresponds to a specific interrupt or exception handler.
the Reset handler is the first function that the CPU will execute, it performs certain actions:
1- copy the .data sections from Flash memory to SRAM memory.
2- initializes the .bss section in SRAM to 0
3- call the main function
To understand better what is happening here, let's understand the boot up proceeder.
When the CPU (specifically the Cortex-M4) starts up, it uses a register called the Program Counter (PC), which always holds the address of the next instruction to be executed by the CPU.
At boot time, the PC is initialized with the value stored at address 0x00000004. However, before this happens, the Cortex-M4 first loads the value from address 0x00000000 into the Stack Pointer (SP). The address 0x00000000 (where the VT is stored) contains the initial SP value, which the CPU uses to set up the stack.
Note: i mentioned before that sometimes the first element of the VT is in some MCU's is reserved, which means that the address of SP is hard coded.
this is done because other MCUs supports extern SRAM, so you can choose which SRAM to use (the internal provided by ST or the external one).
After loading the Stack Pointer, the CPU sets the Program Counter to the value found at address 0x00000004, which corresponds to the Reset Handler. This is the second entry in the Vector Table.
The CPU then begins executing the code at the Reset Handler, which initializes SRAM (init the .data and .bss sections) and eventually calls the main function to start executing your program.
To sum up:
1- Fetch the Stack Pointer: The CPU loads the initial Stack Pointer value from 0x00000000.
2- Load the Reset Handler: The CPU sets the Program Counter to the value at 0x00000004.
3- Execute the Reset Handler: The CPU starts executing the Reset Handler, which eventually leads to the execution of the main function.
now you have a complete picture of what is a startup file, lets write some code.
you can either write it in assembly or in C, i am using C.
so we start by defining the VT, which contains the address of the Handler, we can use the Refrence manual, remember to respect the order and replace the reseved spot with 0.
It depends on the MCU you are using, Since i am using STM32F4RET6, my startup looks like this startup.c
it is decalared as an array:
uint32_t vector_table[] __attribute__((section(".isr_vector"))) = {};the __attribute__((section(".isr_vector"))) is a gcc attribute defines a new section. ( we put the VT to a new section we created)
so wee need to update the linker script to include this section:
SECTIONS {
.text : {
*(.isr_vector)
*(.text)
*(.rodata)
}
}lest's write the definition of each handler
let's start with the Reset_Handler
void Reset_Handler(void); for the other handler, they all follow the same style:
void NMI_Handler(void) __attribute__((weak, alias("Default_Handler")));weak: This is a GCC attribute that tells thelinkerthat this function can be overridden by another function with the same name. If the user defines their ownNMI_Handler, it will replace thisweakdefinition.alias: This attribute creates an alias for the function, in this case, pointing to theDefault_Handler. If anNMI interruptoccurs and theNMI_Handlerisn't defined elsewhere, the CPU will executeDefault_Handler.
Here’s what the Default_Handler looks like:
// defenition
void Default_Handler(void);
// implementation
void Default_Handler(void){
while(1);
}note: the use of a default handler is really important
- we can't implement all the handlers, we leave it for the user to do so
- we need this function to catch it, this is useful when an exception occurs like division by
Owe keep the CPU stuck, or else it will enter unpredicted sate (better be safe than sorry ;) ) .
The definition of the handler are here
Will see later how to implement the Reset_Handler.
In this section, I'll show you how to define the first few elements of the Vector Table (VT). You can complete the rest on your own based on this example:
uint32_t vector_table[] __attribute__((section(".isr_vector"))) = {
0,// in the stM32 NUCLEO-F4RE RF-manual the first elem of the VT is reserved (doesn't have an externel SRAM so the SP is hard-coded )
(uint32_t)Reset_Handler,
(uint32_t)NMI_Handler,
(uint32_t)HardFault_Handler,
(uint32_t)MemManage_Handler,
(uint32_t)BusFault_Handler,
(uint32_t)UsageFault_Handler,
0, // Rerved
// ...
};Note: We cast the function pointers to uint32_t because the functions are of void type, while the array expects uint32_t elements. This ensures the addresses of the handlers are correctly stored in the Vector Table.
this is where all the fun is.
as mentioned before, the Reset_Handler
1- copy the .data sections from Flash memory to SRAM memory.
2- initializes the .bss section in SRAM to 0
3- call the main function
let's start with the first: one copying the .data section:
we need to get the size of the .data section, how we do that from the linker script by utilizing the . (dot) function.
#TODO add a picture explaining the process
we update the linker script to be like this:
SECTIONS
{
.text :
{
*(.isr_vector)
*(.text)
*(.rodata)
. = ALIGN(4) ;
_etext = . ;
}>FLASH
/* data section */
.data :
{
_sdata = . ;
*(.data)
. = ALIGN(4) ;
_edata = . ;
}> SRAM AT> FLASH
/* uninitialized data */
.bss :
{
_sbss = . ;
*(.bss)
. = ALIGN(4) ;
_ebss = .;
}> SRAM
}you can observed that i added these variables:
_etext: Marks the end of the.textsection inFlash. This symbol helps determine where the initialized data begins inFlash._sdata: Indicates thestartof the.datasection inSRAM. This is where the initialized variables will be copied to._edata: Marks theendof the.datasection inSRAM, used to determine the size of the.datasection._sbss: Indicates thestartof the.bsssection inSRAM._ebss: Marks theendof the.bsssection in SRAM
Note: The align is used to align the addresses by 4 bytes. (explained waht aligned is earlier)
Now we obtained the location of each section, we can export them to the startup file using the C keyword extern:
extern uint32_t _etext;
extern uint32_t _sdata;
extern uint32_t _edata;
extern uint32_t _sbss;
extern uint32_t _ebss;Then can start happily implementing the reset handler:
- first lest copy
.datasection from theflashtoSRAM
// copy .data section to sram
uint32_t size = &_edata - &_sdata ; // calculate the size of .data section
uint8_t *pDst = (uint8_t*)&_sdata; //start position of .data in SRAM
uint8_t *pSrc = (uint8_t*)&_etext; //start position of .data in FLASH
for (uint32_t i= 0 ; i < size ; i++ ){
*pDst++=*pSrc++;
}- second initialize the
.bsssection to0.
// init the .bss section to Zero in SRAM
size = &_ebss - &_sbss ; // size of .bss
pDst = (uint8_t*)& _sbss; // the start position of .bss in SRAM
for (uint32_t i= 0 ; i < size ; i++ ){
*pDst++ = 0 ;
}- finally we call
main:
// put it at top
extern void main(void);
// implementation
void Reset_Handler(void) {
// copy .data section to sram
uint32_t size = &_edata - &_sdata ;
uint8_t *pDst = (uint8_t*)&_sdata; //flash
uint8_t *pSrc = (uint8_t*)&_etext;//sram
for (uint32_t i= 0 ; i < size ; i++ ){
*pDst++=*pSrc++;
}
// init the .bss section to Zero in SRAM
size = &_ebss - &_sbss ;
pDst = (uint8_t*)& _sbss;
for (uint32_t i= 0 ; i < size ; i++ ){
*pDst++ = 0 ;
}
// call main
main();
}in this section i talked about how a general overview of interuupts, an sneek peak how stm32f4 boots up and how to write a startup file
#TODO
A toolChain is collection of binaries which allows to:
- compiler : compiles your code (
C/C++,..) to an assembly code. - assembler : transforms the assembly code to object files
.o(binary format) - linker : takes all the object files combines them, to obtain the final executable
.elf. - Build systems: automates the building process of the final executable
make, camke, ... - Libraries: combination of function you use a cross your application (
math.h) - Debuggers : allows you to inspect your code at runtime, step by step.
- Programmer : a utility that uploads the final executable on your target MCU.
- emulator : a tool that emulates you hardware used for testing purposes
For our application, we will be using arm-gcc(compiler, assembler and linker), for the momement we don't need a library,
as the build system, i am using make, the debugger is GDB and for the programmer i am using the open-source tool open-ocd
this is a cross-compiler that compiles the code on host machine (you machine ie: x86_64) for tagret architechure arm.
the arm-gcc also come with these utilities:
- dissect different sections
- disassemble
- extract symbol and size information
- convert executable to other formats, bin,hex
- Provides C standard libraries
Toolchains for arm:
arm-gcc: GNUtoolchain(Free and open-source)armccfrom ARM, (ships withKEIL, code restricted, requires licensing) linux:sudo apt install gcc-arm-none-eabior
wget "https://developer.arm.com/-/media/Files/downloads/gnu-rm/10.3-2021.10/gcc-arm-none-eabi-10.3-2021.10-x86_64-linux.tar.bz2"
tar -jxf gcc-arm-none-eabi-10.3-2021.10-x86_64-linux.tar.bz2
rm gcc-arm-none-eabi-10.3-2021.10-x86_64-linux.tar.bz2
export PATH="/opt/gcc-arm-none-eabi-10.3-2021.10/bin:$PATH"You can put the export command in the .bashrc file
arm-none-eabi-gcc file.c -o file-omeans specifies the output file name-E: stop after the pre-processing stage, output file format.i-c: tells the compiler to compile and assemble theCfile but not link, output file format (`.o)-Sstop after the compilation stage, do not assemble, output file format.S-march=[name]: this specifies the architecture that the assembler will use, (each arch has it's specific assembly language)
-mcpu=[name]: specifies the name of the target arm processor, GCC will use this name to derive the target architecture (march) and the ARM processor type for which to tune for performance (mtune)-mthumbor-marmby default it's set to-marm: select between generating code that executes inARMandThumbstates
Note
- Use
-mthumbif you need smaller code size and are working in an environment where memory is limited, and performance isn't the highest priority. (tells the assembler to use 16-bit long instructions) .- Use
-marmif you need maximum performance, particularly for tasks that require heavy computation, and memory size isn't a primary concern.(tells the assembler to use 32-bit long instructions) .
arm-none-eabi-gcc -c main.c -mthumb -mcpu=cortex-m4 -o main.oOther flags:
-OO: no optimization
-std=gnu11 specifies what C standard to use
As i have said before, make is build system that automates the compilation process of your project
todo
we will understand the diff between the static library and dynamic library, create a static library and include it.
todo
how to use the open-ocd to flash the compiled code
todo
will understand how GDB works and how to use it
todo
how QMU works and how to use it we will look also to renode
todo